Bring a model. Shrink it. Run it.
This is the whole product in one scene: hand RAI a checkpoint, watch the converter pack its weights to 4-bit, see the file shrink by the bit ratio, load it into the RAM you already have — and put your CPU to work.
Everything the usual stack needs, it doesn't.
RAI runs on the CPU you already have. Throw the switch to see both architectures side by side.
GPU-SIDE ROWS DESCRIBE THE GENERIC ALTERNATIVE THIS ENGINE WAS WRITTEN TO AVOID — NOT ANY NAMED PRODUCT.
One pass, end to end. Run it.
Prompt to stream through quantized kernels, GQA, mixture-of-experts routing, a KV cache, and a sampler whose temperature, top-k, top-p and speculative-decode switches are live — the simulation samples for real.
Press RUN PASS and watch one token travel the whole engine.
Drag the model. Watch the RAM.
Parameters, context window, activation width — slide them and the memory envelope moves. Arithmetic from stated quant widths, not measured benchmarks.
DERIVED ESTIMATES FROM STATED QUANT WIDTHS AND CONTEXT — ARITHMETIC, NOT MEASURED PERFORMANCE.
The rules, demonstrated.
Loopback-only serving, weights that never leave the machine, and an accelerator bay that stays empty — press everything.
rai serve publishes Studio and its HTTP API. Choose what it listens on:
> listening on 127.0.0.1 — studio readyWeights never enter a repository and never leave the machine. Run the conversion:
model.q4.raimodel · local diskCLOUD SLOT — NO UPLOAD PATHNo CUDA. No ROCm. No Metal. Lift the dust cover: