Programmable AI inference by CantorAI

Models are programs.
Compile them for the hardware you ship.

Garnet captures XLang3 Tensor Expressions from backend-neutral xModel programs, then compiles them into hardware-specialized execution through TensorRT, OpenVINO, native CPU algorithms, or custom kernels.

LLMVLMASRTTS
GARNET COMPILER / RUNTIMEENGINE READY
XMODEL IR PAGED KV NATIVE I/O
BackendTensorRT
PrecisionBF16 / INT4
ExecutionNative
01

Models are programs

Readable XLang3 .py source describes symbolic tensor computation that Garnet can inspect and compile.

02

Backend-neutral

xModel semantics stay separate from TensorRT, OpenVINO, native CPU, and custom hardware policy.

03

Hardware-specialized

Garnet lowers graphs, operators, precision, memory behavior, and kernels for the machine you deploy.

Open xModel programs

Four modalities.
One programmable runtime.

Inspect the model programs, manifests, and backend profiles in the public repository. Checkpoints and vendor runtimes remain external.

VLM · IMAGE + TEXT SOURCE INCLUDED

Qwen3-VL 2B

Vision-language inference with native image preprocessing, MRoPE, continuous batching graphs, and paged KV.

TensorRT BF16NVIDIA GPUJPEG input
View xModel source →
ASR · SPEECH TO TEXT SOURCE INCLUDED

Qwen3-ASR 0.6B

Native WAV ingest, resampling, log-mel transform, audio encoding, and text decoding in one compiled path.

TensorRT BF1616 kHz inputNative frontend
View xModel source →
TTS · TEXT TO SPEECH SOURCE INCLUDED

Qwen3-TTS 0.6B

CustomVoice generation with native codec reconstruction and 24 kHz waveform output.

TensorRT BF1624 kHz outputSpeaker selection
View xModel source →

Garnet Runtime

Model code is not
a deployment plan.

Garnet executes Python-compatible xModel source in XLang3, captures an inspectable Tensor Expression graph, and lowers it into execution specialized for the selected backend and processor architecture.

  • ✓ Native C++ execution—no Python process at inference
  • ✓ TensorRT, OpenVINO, native CPU, and custom-kernel execution
  • ✓ Paged KV memory, fixed decode buckets, and engine caching
  • ✓ Validated model manifests and application-facing APIs
Request a technical briefing →
COMPILED MODEL PIPELINE
01
Selected modelSafetensors + configuration
Qwen3
02
xModel programXLang3 Tensor Expressions
.py
03
Target profileBackend · precision · shape
JSON
NVIDIATensorRTGPU engines
INTELOpenVINOCPU / GPU

Build Garnet from source

Inspect the graph.
Compile for your hardware.

Clone Garnet beside XLang3, build the native package, validate graph capture with a weight-free xModel, and then bring the checkpoint and vendor runtime for your target backend.

Windows x64 documented pathLinux support in CMakeApache-2.0
DEVELOPER FLOW4 STEPS
  1. 01
    Clone the source

    Place Garnet and XLang3 in sibling directories.

  2. 02
    Build the runtime

    Configure the C++ toolchain and the backend SDKs required for your target.

  3. 03
    Capture an xModel

    Execute Python-compatible model source with symbolic tensors and inspect the graph.

  4. 04
    Compile and run

    Lower the graph for TensorRT, OpenVINO, native CPU, or custom hardware execution.

Explore integration examples →

Measured, not imagined

Performance with context.

Every number names the model, hardware, precision, and workload behind it.

NVIDIA · THROUGHPUT
527.9 tok/s

Qwen3-1.7B · batch 4

Masked continuous decode on RTX 4080 using official BF16 weights and GPU-resident paged KV storage.

INTEL · CPU INFERENCE
3.33×

Garnet vs. Ollama

31.87 vs. 9.56 decode tok/s on an i9-14900K, comparing Garnet INT4 with the tested Ollama package.

GarnetOllama
OUR STANDARD

Speed is only useful when the output stays useful.

Garnet validation records correctness, response quality, cold-load behavior, cache state, and unsupported hardware—not just the best token counter.

Ask for benchmark details →

Internal engineering measurements. Results vary by model, prompt, driver, operating system, and engine cache state. Comparisons are not precision-equivalent unless explicitly stated.

Custom model engineering

Your model.
Your hardware.
A measured path to production.

Have a model or device target outside the catalog? CantorAI can evaluate the graph, memory budget, precision strategy, native I/O path, and application packaging.

Start an optimization review →

Ready to program closer to the hardware?

Read the model. Inspect the graph. Own the execution.

View Garnet on GitHubGet started