A CPU-only LLM inference engine written in Rust. RAI runs 4-bit quantized
language models with hand-written AVX2 kernels — no GPU, no CUDA, no Python
runtime, no PyTorch, no GGML, no BLAS. Load a .raimodel file and generate
text on any supported x86-64 machine.
Built by ClassEve. Licensed under Apache-2.0.
Official repository. This is the only official repository for RAI. ClassEve's complete list of official accounts is at classeve.com/official. The GitHub account
github.com/ClassEveis an unrelated third party, not affiliated with ClassEve.
Measured performance
Measured on a consumer-grade laptop CPU (4 cores / 8 threads), 2026-08-09, RAI 0.2.0. Full method, roofline, and the results that came out negative are in BENCHMARKS.md.
Converting a checkpoint — rai convert streams .safetensors, so peak
memory does not grow with the model:
export_rtn.py (PyTorch) |
rai convert |
|
|---|---|---|
| TinyLlama-1.1B | 188.8 s, 4,981 MB RAM | 7.6 s, 22.9 MB RAM |
| Zephyr-7B | needs ~29 GB — will not run | 82.6 s, 26.3 MB RAM |
Both produce byte-identical output. A 7B model converts on a 16 GB laptop.
Running a model — greedy decoding, same machine:
| TinyLlama-1.1B (619 MB) | Zephyr-7B (3.9 GB) | |
|---|---|---|
| Decode | 21.8 tok/s | 2.96 tok/s |
| Peak RSS | 629 MB | ~4.0 GB |
| Load (warm) | 0.33 s | — |
HuggingFace transformers fp32 runs the same TinyLlama checkpoint at 4.3 tok/s
on this machine, so RAI is 5.1× faster in 1/7th the memory. The batched-GEMM
rewrite made prefill ~1.3× faster, and prompt-lookup decoding adds
1.12–1.20× on context-quoting workloads (off by default; it is slower on
original prose, and BENCHMARKS.md says by how much).
The decode figure is a quiet-machine number, re-verified 2026-08-18 at
22.1–23.0 tok/s on the shipped x86-64-v2 baseline; the same binary under
background load on this 4-core laptop measured as low as 2 tok/s.
BENCHMARKS.md has the full record and the reconciliation
with the 0.2.2 changelog's loaded-machine figures.
Quickstart
# 1. Build (or take a release archive — see INSTALL.md) cargo build --workspace --release --locked # 2. Convert a checkpoint — no Python required # (-o sets the output name; convert otherwise preserves the checkpoint's case) rai convert /path/to/Qwen2.5-0.5B-Instruct -o qwen2.5-0.5b-instruct-q4.raimodel # 3. Generate rai run qwen2.5-0.5b-instruct-q4.raimodel \ --chat-template chatml \ --prompt "Explain photosynthesis in simple terms." \ --max-tokens 64 # 4. Or open the local chat interface rai serve qwen2.5-0.5b-instruct-q4.raimodel # What is on this machine? rai models .
Conversion writes tokenizer.json beside the model, and rai run picks it up
automatically. Instruction-tuned models need the chat template they were
trained on, or they emit end-of-sequence immediately and print nothing:
chatml for Qwen, llama3 for Llama-3, zephyr for TinyLlama-Chat.
Prebuilt binaries are in INSTALL.md; full conversion options,
including the calibrated GPTQ path, are in
docs/INSTALL.md. The double-click
launchers in launchers/ start the local interface without a terminal.
Which models work
RAI runs the Llama-family decoder and the capabilities layered on top of it. Anything it cannot represent is refused at conversion time, by name, before a file is written — never silently mis-converted.
| Runs | Llama-2 / 3 / 3.1 / 3.2, Mistral-7B, Qwen2 / 2.5 / 3 (dense), Gemma / Gemma2 / Gemma3, Phi-3 / 3.5, OLMo2, mixture-of-experts models (OLMoE, Mixtral, Qwen3-MoE), TinyLlama, SmolLM / SmolLM2, and fine-tunes of all of them |
| Refused | Models with a shared expert running alongside the routed ones; rope_scaling other than default or llama3 (yarn, linear, dynamic); an lm_head bias; and module trees that are not Llama-shaped — Falcon, GPT-NeoX, MPT, GPT-2, Phi-2. Each refusal names the blocker and a model that works. |
Point rai convert at your folder, or press Check in Studio: both run the
same preflight, refuse before writing anything, name the blocker, and name a
model that does work. What to use instead of a refused checkpoint, and what has
to be in the folder, are in docs/MODELS.md.
Every family above was converted and generated coherent text on the
machine in the benchmarks: Qwen2.5-0.5B, Qwen3-0.6B, Llama-3.2-1B, gemma-2b-it,
gemma-2-2b-it, gemma-3-1b-it, Phi-3-mini-4k-instruct, OLMo-2-1B-Instruct and
OLMoE-1B-7B-Instruct (6.9B parameters, 64 experts, 8 routed per token).
A folder holding a pytorch_model.bin rather than .safetensors is not read —
that is a Python pickle, and loading one executes whatever code the file
carries. The error names the file and hands you the one-line command that
converts it.
Why
Most LLM inference stacks assume a GPU, a CUDA toolchain, or a heavyweight ML
framework. RAI takes the opposite bet: an auditable Rust workspace whose
inference matrix kernels are hand-written and whose runtime dependency tree is
captured in Cargo.lock.
- CPU-only by design. AVX2 + FMA + F16C accelerate inference on compatible x86-64 CPUs; scalar fallbacks exist for other instruction sets.
- 4-bit weights, dequantized in registers. Weights stay packed in memory; unpacking happens on the fly inside the GEMM inner loop. No fp32 weight copy ever exists in RAM.
- One flat model file. The
.raimodelformat is a single binary blob with a 128-byte header. The loader validates its structure after one heap read and then exposes borrowed views over the in-memory sections. - A lean library.
--no-default-featuresbuilds the inference library — format reader, kernels, model, sampling, speculative decoding — againsthalf,rayon,anyhow, andrandalone. The CLI, tokenizer, and chat server sit behind the default-onclifeature. - Speculative decoding. Draft-model and prompt-lookup modes. Both accept or reject each drafted token against the target model's own sampled distribution, with a correction sample on rejection; a seeded statistical smoke test guards that path.
- Local serving. An HTTP chat server with a built-in web UI, plus a REST + MCP server so agentic tools (e.g. Claude Desktop, Claude Code) can use RAI as a tool backend.
Workspace layout
| Crate | Purpose |
|---|---|
rai-infer |
The inference engine: .raimodel loader and writer, AVX2 W4A8 GEMM kernels, transformer layers (RMSNorm, RoPE, GQA, SwiGLU/GeGLU), KV cache, sampling, speculative decoding, and the rai binary (behind the default-on cli feature; --no-default-features builds the lean library) |
rai-compress |
Quantization and compression research toolkit. Its Rust GPTQ implementation is independent of the Python .raimodel export pipeline; RC/HRC/SAC report modeled sizes and serialize no artifact. Nothing here is on the inference path. |
RAI is rai-infer and rai-compress. The three crates below are a separate
memory/reasoning service that only shares this workspace. They are not part of
the RAI product, are not published to crates.io (0.1.0 yanked 2026-08-13,
publish = false set), and rai-server imports rai-infer zero times — it
cannot run a model.
These three are absent from every release archive, and a release-time gate fails the build if one of them reaches an archive again.
| Not part of RAI | Purpose |
|---|---|
rai-server |
REST + MCP server for the memory/reasoning layer |
rai-core |
Memory, embedding and reasoning primitives used by rai-server |
rem-nra |
Resonance-memory backend used by rai-core |
The inference and memory-service paths are separate. rai run and rai serve
load .raimodel files through rai-infer; rai-server does not run those
models. Instead, its REST/MCP adapters call rai-core, which obtains an
embedding from the configured provider and stores/queries state through
rem-nra. AppState serializes REST stores and opted-in MCP stores to
RAI_DATA_PATH.
Requirements
| Requirement | Details |
|---|---|
| Rust | 1.87+; the repository pins 1.95.0, edition 2021 |
| CPU | x86-64 with AVX2, FMA, and F16C for optimized paths; scalar fallbacks otherwise |
| OS | Linux, Windows, or macOS |
| GPU at runtime | Not required |
| Python | Calibrated (GPTQ) export and draft-model preparation only. rai convert does round-to-nearest conversion without it. |
.cargo/config.tomlpins the x86-64 build to the x86-64-v2 baseline — the same floor the release archives use — so a binary you build here runs on any x86-64 machine you copy it to. This costs nothing measurable: the AVX2, FMA, and F16C kernels are chosen at runtime, not by the compile-time baseline. aarch64 (Apple Silicon, ARM servers) is left at the toolchain default. To tune a build to the machine in front of you — for a local benchmark, never for a binary you hand to anyone else —RUSTFLAGS="-C target-cpu=native" cargo build --release --locked; that binary dies with SIGILL on any older CPU.
Build
cargo build --workspace --release --locked
That produces rai, plus the deprecated rai-convert, rai-generate, and
rai-chat wrappers — and, because --workspace builds everything in the
tree, the non-product rai-server too. Release archives are built from
--package classeve-rai-infer alone. See installation
for source installs and the container. No container image is published; the
Dockerfile builds the rai CLI by default, and
docker build --target server . builds the non-product memory service
instead.
Development checks:
cargo fmt --all -- --check cargo clippy --workspace --all-targets --locked -- -D warnings cargo test --workspace --all-targets --locked cargo test --workspace --doc --locked
Converting a model
RAI runs models in its own .raimodel format. Two paths produce it:
# Round-to-nearest — no Python, no torch rai convert /path/to/Mistral-7B-Instruct-v0.3 --max-context 4096 # GPTQ, calibrated — needs the Python environment and a calibration corpus python3 rai-infer/scripts/export_raimodel.py \ --model /path/to/SmolLM-135M \ --output smollm-135m-q4.raimodel
Both write the model file plus a tokenizer.json alongside it, and both refuse
architectures the format cannot represent. They do not cover the same models:
rai convert writes container v2 and handles Qwen2/2.5, Qwen3 (dense and
MoE), Llama-3.1/3.2, Gemma 1/2/3, OLMo2, Phi-3, and Mixtral, while the Python
exporters write container v1 and refuse all of those.
Check docs/MODELS.md before downloading a checkpoint;
docs/INSTALL.md has every flag.
HuggingFace model and dataset revisions are not pinned by the Python exporters,
so record the exact revisions and arguments for reproducible work.
Running
Generate text
rai run smollm-135m-q4.raimodel \
--prompt "The future of computing is" \
--max-tokens 64 \
--temperature 0.7--tokenizer defaults to the tokenizer.json written beside the model at
conversion time; pass it explicitly only for a model that was moved away from
its tokenizer.
Sampling controls: --temperature, --top-k, --top-p,
--repetition-penalty, --seed. Test-time compute ("pondering") strategies
exist behind --ponder-strategy cfg|ensemble|cfg-ensemble|adaptive with
--guidance-scale, --ensemble-n, --noise-sigma, --entropy-threshold;
they multiply the forward passes per token and no measurement in this
repository shows a quality win from them — the module doc in
rai-infer/src/ponder.rs says exactly what each one computes. Leave them off.
Set RAYON_NUM_THREADS to cap the inference worker count used by rai run and
rai serve; it defaults to Rayon's own choice.
Speculative decoding
Two modes, mutually exclusive. Both verify against the target model and are
gated on exact sampling (--top-k 0 --top-p 1 --repetition-penalty 1), so
verification uses the distribution the target actually produced.
# Draft-model speculation: a small model proposes, the big model verifies rai run mistral-7b-q4.raimodel \ --draft mistral-draft-100m-q4.raimodel \ --draft-k 6 \ --top-k 0 --top-p 1 --repetition-penalty 1 \ --prompt "Explain speculative decoding in one paragraph." # Prompt-lookup: the draft is copied from the context, so there is no draft # model and no draft forward pass rai run tinyllama-q4.raimodel \ --lookup-k 2 \ --top-k 0 --top-p 1 --repetition-penalty 1 \ --prompt "Summarise the passage above."
Draft and target must share a tokenizer. Prompt-lookup is off by default: it is
a gain only when the output reuses the context, and BENCHMARKS.md records both
the gain and the loss. rai-infer/scripts/train_draft.py is an experimental
distillation helper for compatible Mistral-family teachers; its throughput
projections are not release benchmarks.
Chat over HTTP
rai serve smollm-135m-q4.raimodel --port 8090
Open http://localhost:8090 for the built-in web UI, or POST to /api/chat
(JSON) for programmatic access — either a single message or a messages
array carrying the whole conversation, oldest first; the server keeps no chat
state of its own, and the web UI replays the visible thread the same way.
--chat-template auto|none|few-shot|mistral|llama3|chatml|zephyr|phi3|gemma
selects prompt formatting. The chat server binds to 127.0.0.1 only and
limits request bodies to 64 KiB.
The memory service is not in the archive
The rai-server REST/MCP memory prototype builds from this workspace but is
not part of RAI and is not in any release archive — see
Workspace layout. It runs models zero times, and it is the
only thing here that can open an outbound connection. If you want to build and
run it from source, docs/OPERATIONS.md is its runbook.
The .raimodel format
A single flat, little-endian binary file:
┌─────────────────────────────────────────────┐
│ Header (64 bytes at v1, 128 at v2) │ magic "RAIM", version,
│ │ architecture hyperparameters,
│ │ quantization config
├─────────────────────────────────────────────┤
│ Section index table (16 bytes per section) │ offset + size per section
├─────────────────────────────────────────────┤
│ Section 0: embedding table (8-bit) │
│ Sections 1..N: transformer layers (4-bit) │ per linear: dims, f16
│ Section N+1: final RMSNorm (f32) │ scale/zero per group,
│ Section N+2: lm_head (4-bit, when untied) │ nibble-packed codes
└─────────────────────────────────────────────┘
Each layer section holds seven 4-bit projections — q, k, v, o, gate,
up, down — then any f32 bias vectors the header's bias_mask declares, then
two f32 RMSNorm weight vectors. Linear weights carry per-group (128-column by
default) f16 scale/zero parameters and are packed two codes per byte in the
layout the AVX2 kernels consume. The embedding table is 8-bit quantized.
Scale/zero parameters are round-tripped through f16 at export time so the reader
and the exporter use the same stored values.
Version 1 holds the architecture dimensions, rope_theta, and norm_eps.
Version 2 extends the header to 128 bytes and adds the activation code, the
llama3 RoPE rescaling parameters, the bias mask, and the embedding scale — the
four fields that brought Qwen2/2.5, Llama-3.1/3.2, and Gemma inside the
supported set. A converter emits v1 whenever none of them is needed, so files
that converted before v2 still produce the same bytes. There is no field for
anything beyond that, which is what makes the compatibility list in
docs/MODELS.md what it is.
The loader performs one read of the whole file into heap memory, validates the
header and section bounds, and then hands out borrowed slices into that buffer.
This avoids additional copies between parsed sections, but loading still copies
the file from storage into the process heap. No retained benchmark establishes
a general performance advantage over mmap.
More measurements
The headline table is at the top of this file. Beyond it, BENCHMARKS.md records
the roofline analysis (decode reaches 45% of this machine's 26.4 GB/s memory
ceiling), the batched-GEMM prefill rewrite, and the per-tensor quantization
error, which ranged from 6.4e-06 (q_proj) to 6.2e-09 (the 8-bit embedding).
It also records what was measured and rejected, so nobody re-derives it: self-speculative early exit reached 0.4% draft acceptance and ran roughly 15× slower than plain decoding, so it is not in the CLI — the library implementation remains for use with a trained exit head. An older author-run measurement reported ~195 tokens/s for SmolLM-135M, but its raw output and environment were not retained; it is kept as historical context, not as release evidence.
See BENCHMARKS.md for the full method and the quantization-quality comparison.
Project status
RAI is pre-1.0 and interfaces may change. What that means concretely:
- One end-to-end conversion and decode run is measured (TinyLlama-1.1B, BENCHMARKS.md). The rest of the supported list is verified by architecture and, for Qwen2.5-0.5B, Qwen3-0.6B, Llama-3.2-1B, gemma-2b-it, gemma-2-2b-it, gemma-3-1b-it, Phi-3-mini-4k, OLMo-2-1B and OLMoE-1B-7B, by a coherent generation — not by a retained qualification matrix.
- Architecture coverage is enforced at conversion time, and what remains
unsupported is unsupported for something a checkpoint contains rather than
for its name: a shared expert, a
rope_scalingscheme the kernels do not implement, anlm_headbias, or a module tree that is not Llama-shaped. docs/MODELS.md names each one and what to use instead. - 4-bit conversion costs real quality, and the number is now measured. On
wikitext-2,
rai convert's round-to-nearest 4-bit raises perplexity by 13% to 30% over the fp16 checkpoint, across Qwen2.5-0.5B/1.5B/3B and SmolLM2-1.7B, and at--temperature 0the 4-bit model emits a different token than fp16 between one time in four and one time in five. llama.cpp's Q4_K_M costs three to eight times less on the same corpus, because it is not uniformly 4-bit — it spends 6 bits onffn_downandattn_v, where RAI spends 4 on everything. RAI's files are correspondingly smaller. Method, machine, corpus hash and the full tables are in BENCHMARKS.md; reproduce any of it withrai perplexity. - The optimized paths are x86-64 (AVX2 + FMA + F16C). On ARM — including
Apple Silicon, which ships a native
aarch64-apple-darwinarchive — the GEMM falls back to a scalar path, so it runs correctly but far slower than the measured x86-64 figures above. NEON kernels are future work, and no ARM throughput is claimed here because none has been measured. - The memory/reasoning service and the RC/HRC/SAC compression paths are research prototypes. Their confidence, crowding, and modeled-size outputs are not validated guarantees.
Issues and PRs are welcome. See SECURITY.md, SUPPORT.md, and CONTRIBUTING.md.
About
Built and maintained by ClassEve — engineering for AI agents and developer tooling. Project page: classeve.com/public/rai.
License
Apache License 2.0 — see LICENSE. Copyright 2025-2026 ClassEve.