Models are programs
Readable XLang3 .py source describes symbolic tensor computation that Garnet can inspect and compile.
Programmable AI inference by CantorAI
Garnet captures XLang3 Tensor Expressions from backend-neutral xModel programs, then compiles them into hardware-specialized execution through TensorRT, OpenVINO, native CPU algorithms, or custom kernels.
Readable XLang3 .py source describes symbolic tensor computation that Garnet can inspect and compile.
xModel semantics stay separate from TensorRT, OpenVINO, native CPU, and custom hardware policy.
Garnet lowers graphs, operators, precision, memory behavior, and kernels for the machine you deploy.
Open xModel programs
Inspect the model programs, manifests, and backend profiles in the public repository. Checkpoints and vendor runtimes remain external.
Compact language generation with native chat templating, paged KV cache, and optimized decode.
Vision-language inference with native image preprocessing, MRoPE, continuous batching graphs, and paged KV.
Native WAV ingest, resampling, log-mel transform, audio encoding, and text decoding in one compiled path.
CustomVoice generation with native codec reconstruction and 24 kHz waveform output.
Higher-capacity CustomVoice generation with speaker selection, instruction control, and native 24 kHz codec reconstruction.
Garnet Runtime
Garnet executes Python-compatible xModel source in XLang3, captures an inspectable Tensor Expression graph, and lowers it into execution specialized for the selected backend and processor architecture.
Qwen3.pyJSONBuild Garnet from source
Clone Garnet beside XLang3, build the native package, validate graph capture with a weight-free xModel, and then bring the checkpoint and vendor runtime for your target backend.
Place Garnet and XLang3 in sibling directories.
Configure the C++ toolchain and the backend SDKs required for your target.
Execute Python-compatible model source with symbolic tensors and inspect the graph.
Lower the graph for TensorRT, OpenVINO, native CPU, or custom hardware execution.
Measured, not imagined
Every number names the model, hardware, precision, and workload behind it.
Masked continuous decode on RTX 4080 using official BF16 weights and GPU-resident paged KV storage.
31.87 vs. 9.56 decode tok/s on an i9-14900K, comparing Garnet INT4 with the tested Ollama package.
Garnet validation records correctness, response quality, cold-load behavior, cache state, and unsupported hardware—not just the best token counter.
Ask for benchmark details →Internal engineering measurements. Results vary by model, prompt, driver, operating system, and engine cache state. Comparisons are not precision-equivalent unless explicitly stated.
Custom model engineering
Have a model or device target outside the catalog? CantorAI can evaluate the graph, memory budget, precision strategy, native I/O path, and application packaging.
Start an optimization review →Model, workload, hardware, latency, memory, and quality targets.
Graph lowering, quantization, kernels, scheduling, and native I/O.
Correctness, quality, throughput, latency, startup, and recovery.
A versioned model package ready for your application.
Ready to program closer to the hardware?