Speculative inference for heterogeneous machines and consumer GPUs.
Custom kernels, speculative prefill and decoding, tuned for each model and hardware target.
Inference Engine Optimizations
| Optimization | Measured setup | Result |
|---|---|---|
| DFlash2 | Qwen 3.8 27B on one R9700 | 208.1 tok/s average, 227.8 tok/s peak |
| DSpark | DeepSeek V4 on Strix Halo, native top-6 | 32.7 tok/s high-acceptance median; 27.9 tok/s mixed-eval average |
| PFlash + KVFlash | Laguna XS 2.1 33B at 256K on RTX 3090 | 6.1× prefill, 411 s to 67.3 s |
| Luce Spark | Laguna XS.2 33B on RTX 3090 | ~100 tok/s in 14.6 GiB |
| KVFlash | Laguna XS 2.1 33B at 256K on RTX 3090 | 152.3 tok/s with an 8K pool |
| Heterogeneous execution | DeepSeek V4 on R9700 + Strix Halo | 86 tok/s decode; 788 tok/s prefill at 2K |
| Paged attention + continuous batching | Qwen 3.8 27B + DFlash2 on R9700; DeepSeek V4 Flash AR on Strix Halo | 300.9 tok/s total at 5 clients (Qwen); 48.4 tok/s output-window at 4 clients (DeepSeek) |
| Megakernel | Qwen 3.5 0.8B on RTX 3090 | 413 tok/s, 1.87 tok/J |
| Vision (image input) | Qwen 3.8 27B vision on R9700 + DeepSeek V4 Flash Vision on Strix Halo, together | 3.2× the image-question throughput of a DGX Spark (58 vs 18 a minute); 2.4× faster at 8 users with the same model file |
Supported Models and Drafters
Model links open the exact weights used by the measured setup. Drafter links open the published quant, or the source checkpoint when conversion is required.
| Model and optimization | Phase | Speedup |
|---|---|---|
| Qwen 3.5 0.8B BF16 + Megakernel | Prefill + decode | 1.9× prefill; 1.55× decode |
| Qwen 3.8 27B UD-IQ4_XS + DFlash2 source, converted to Q8_0, on R9700 | Decode | 6.4× vs Lucebox AR; 3.8× vs llama.cpp with the same drafter |
| Laguna XS 2.1 33B Q4_K_M + PFlash/KVFlash with Qwen3 0.6B Q8_0 | Prefill | 6.1×, 411 s to 67.3 s at 256K |
| Laguna XS 2.1 33B Q4_K_M + DFlash Q4 drafter | Decode | 1.7× at 256K |
| Gemma 4 26B-A4B Q4_K_M + DFlash Q8_0 drafter | Decode | 1.31× |
| Gemma 4 31B IT Q4_K_M + DFlash Q8_0 drafter | Decode | 3.2× |
| DeepSeek V4 Flash ROCmFPX MIX Strix + DSpark Q4RMFP4 drafter | Decode | 42 tok/s at 8K and 39 tok/s on code and math with the plain launch (PR #729) |
| Ling 3.0 Flash 124B-A5.1B Q4_K_M | Decode | 34.6 tok/s median AR on DGX Spark |
| Qwen 3.8 27B IQ4_XS + Q8_0 vision projector + DFlash2 drafter, on R9700 (image input) | Image questions | 2.4× vs llama.cpp on DGX Spark at 8 users (10.5 s vs 25.4 s); 2.1× for one user |
| DeepSeek V4 Flash Vision ROCmFPX MIX Strix + BF16 vision projector, encoder on R9700, 4 users at once (image input) | Image questions | 1.5× sooner first token with 16 images when the R9700 encodes |
Tested Machines (GPU/APU)
The engine is not tied to one reference card. NVIDIA architectures are selected by CMake; HIP builds should target the device's exact gfx architecture.
| Architecture | Hardware | Runtime | Details | |
|---|---|---|---|---|
![]() |
RDNA4 gfx1201 |
Radeon AI PRO R9700 | ROCm 7.2 | Qwen 3.8 R9700 quick start |
![]() |
RDNA3.5 gfx1151 |
Ryzen AI MAX+ 395 / Strix Halo | ROCm 7.2 | DeepSeek V4 Strix profile |
![]() |
RDNA3 gfx1100 |
Radeon RX 7900 XT / XTX | ROCm 6+ | DeepSeek V4 dual AMD profile |
![]() |
Ampere sm_86 |
RTX 3090 | CUDA 12+ | Qwen 3.8 NVLink result and Megakernel results |
![]() |
Blackwell sm_120 |
RTX 5090 | CUDA 12.8+ | Qwen 3.8 single-GPU result |
![]() |
Blackwell sm_121 |
DGX Spark / GB10 | CUDA 12.9 | Qwen 3.5 NVFP4 results |
![]() |
Ada sm_89 |
RTX 4090 | CUDA 12+ | Linux and WSL2 community runs |
![]() |
Turing sm_75 |
RTX 2080 Ti | CUDA 12.0 | DFlash results |
![]() |
Volta sm_70, Pascal sm_61 |
V100, P40 | CUDA 12.0 | CUDA quick start |
| Not pictured | Blackwell sm_110 |
Jetson AGX Thor | CUDA 13.0 | Thor quick start |
Single-device results
| Hardware | Model | Measured result |
|---|---|---|
| R9700 | Qwen 3.8 27B UD-IQ4_XS + DFlash2 source | 208.1 tok/s HumanEval average; 227.8 tok/s best request |
| Strix Halo | DeepSeek V4 ROCmFPX MIX Strix + DSpark Q4RMFP4 | 42 tok/s decode and 320 tok/s prefill at 8K, 36 tok/s at 123K, 39 tok/s on code and math, 25 tok/s on prose, all six routed experts, plain launch (PR #729) |
| RTX 5090 | Qwen 3.8 27B | 110.6 tok/s for a 26,758-token prompt and 1,024-token continuation (PR #637) |
Heterogeneous and parallel results
| Hardware | Configuration | Measured result |
|---|---|---|
| 2x RTX 3090 + NVLink | Qwen 3.8 target tensor parallel + DFlash2 | 79.7 tok/s, 2.16× autoregressive decode (PR #637) |
| RX 7900 XT + Strix Halo | DeepSeek V4 with all six experts + DSpark verification width 4 | 45.0 to 47.7 tok/s decode; 111.2 tok/s prefill at 132,981 tokens (PR #604) |
| R9700 + Strix Halo | DeepSeek V4 across both AMD devices | 86 tok/s decode; 788 tok/s prefill at 2K |
These runs use different prompts, quantizations, and inference policies. They show which configurations work; they are not a cross-hardware ranking.
Recommended Setups
See Recommended server setups for the model and hardware matrix, including single-GPU and mixed-GPU profiles.
The DS4 guide also documents the Strix long-context sparse-verifier profile and Qwen3-0.6B PFlash integration. PFlash is lossy prompt compression; keep it off for exact-retrieval and matched true-context benchmarks.
Client Harnesses
harness/ runs Lucebox through popular coding clients and checks server compatibility.
|
Set the server binary and model paths, then run a launcher:
LUCE_SERVER_BIN=server/build/luce_server \ LUCE_TARGET=server/models/Qwen3.8-27B-UD-IQ4_XS.gguf \ LUCE_DRAFT=server/models/draft/qwen38-dflash2-q8_0.gguf \ MAX_CTX=32768 \ harness/clients/run_codex.sh
See the harness guide for setup, no-draft targets, and benchmarks.
Quick Start With Docker
Prebuilt images on GHCR track main. Mount the weights and serve the OpenAI-compatible API on :8000.
Put the target in |
|
Run the image for your GPU:
# NVIDIA docker run --rm --gpus all -p 8000:8080 \ -v "$PWD/server/models:/opt/lucebox-hub/server/models" \ ghcr.io/luce-org/lucebox-hub:cuda12 # AMD docker run --rm --device /dev/kfd --device /dev/dri \ --group-add video --group-add render --security-opt seccomp=unconfined \ -p 8000:8080 -v "$PWD/server/models:/opt/lucebox-hub/server/models" \ ghcr.io/luce-org/lucebox-hub:rocm
The container picks the GPU the model fits on (a discrete card before an integrated one, else the largest) and sizes the context from that GPU's memory. serve is the default command, so luce_server flags can follow the image name directly (or serve); they replace the values the container would pass. --target-device, --max-ctx and --profile all work:
# Show the GPUs, the model, and the device auto placement would use docker run --rm <gpu flags> -v ... ghcr.io/luce-org/lucebox-hub:rocm devices # DeepSeek V4 on Strix Halo with its qualified profile (DSpark drafter in models/draft/) docker run --rm <gpu flags> -p 8000:8080 -v ... ghcr.io/luce-org/lucebox-hub:rocm --profile ds4-strix # Pin a device docker run --rm <gpu flags> -p 8000:8080 -v ... ghcr.io/luce-org/lucebox-hub:rocm --target-device hip:1
Environment variables such as LUCE_TARGET, LUCE_TARGET_DEVICE, LUCE_MAX_CTX and LUCE_ARGS cover the same settings for compose files; see the header of server/scripts/entrypoint.sh.
Run the Server
This quick start runs the R9700 profile above. The complete flag reference is in the server guide.
# build (ROCm 7.2+, RDNA4) git clone --recurse-submodules https://github.com/Luce-Org/lucebox.git cd lucebox cmake -S server -B server/build-hip -G Ninja \ -DCMAKE_BUILD_TYPE=Release \ -DCMAKE_HIP_COMPILER=/opt/rocm/lib/llvm/bin/clang++ \ -DLUCE_GPU_BACKEND=hip \ -DLUCE_HIP_ARCHITECTURES=gfx1201 \ -DGGML_HIP_MMQ_MFMA=ON \ -DGGML_HIP_NO_VMM=ON cmake --build server/build-hip --target luce_server -j"$(nproc)" # target and DFlash2 drafter mkdir -p models huggingface-cli download unsloth/Qwen3.8-27B-GGUF \ Qwen3.8-27B-UD-IQ4_XS.gguf --local-dir models huggingface-cli download incoai/Qwen3.8-27B-DFlash2 --local-dir models/dflash2 python server/scripts/convert_dflash_to_gguf.py \ models/dflash2/model.safetensors models/qwen38-dflash2-f16.gguf python server/scripts/quantize_dflash_draft.py \ models/qwen38-dflash2-f16.gguf models/qwen38-dflash2-q8_0.gguf --scheme q8_0 # launch the measured profile ./server/build-hip/luce_server models/Qwen3.8-27B-UD-IQ4_XS.gguf \ --draft models/qwen38-dflash2-q8_0.gguf \ --draft-block-size 16 --max-ctx 131072 \ --cache-type-k q8_0 --cache-type-v q8_0 \ --port 8216 curl -s http://127.0.0.1:8216/v1/chat/completions \ -H 'Content-Type: application/json' \ -d '{"messages":[{"role":"user","content":"Write a Python LRU cache."}], "max_tokens":256,"temperature":0}'
To serve up to N concurrent requests, use this launch command with the same Qwen target and DFlash2 drafter. Set N to the desired concurrency (5 below). Qwen automatically sizes the shared KV pool from available GPU memory.
N=5
./server/build-hip/luce_server models/Qwen3.8-27B-UD-IQ4_XS.gguf \
--draft models/qwen38-dflash2-q8_0.gguf \
--draft-block-size 16 --max-ctx 16384 \
--paged-attention --max-concurrency "$N" \
--cache-type-k q8_0 --cache-type-v q8_0 \
--port 8216See Continuous batching in Lucebox for Qwen and DeepSeek V4 Flash results, latency measurements, and launch settings.
Documentation
| Topic | Guide |
|---|---|
| Recommended model and hardware profiles | Recommended setups |
| Runtime parameters | Server parameter reference |
| OpenAI Chat Completions, Responses, and Anthropic Messages | API reference |
| CUDA, HIP, and mixed-device placement | Mixed-backend guide |
| DeepSeek V4 single-device and heterogeneous profiles | DeepSeek V4 guide |
| Image input (Qwen3.8, DeepSeek V4 Flash Vision) | Image input guide |
| Environment variables | Environment reference |
| Server internals | Architecture |
| Client integration and qualification | Harness guide |
| Server engine components | Engine components |
Benchmarks stay with each implementation: DFlash, PFlash, Spark, KVFlash, and Megakernel.
Tutorials
Video tutorials for each optimization and the harness setup.
| Luce Spark ▶ YouTube |
Luce DFlash ▶ YouTube |
Luce Turboquant ▶ YouTube |
| OpenClaw harness setup ▶ YouTube |
Luce PFlash ▶ YouTube |
Luce Megakernel ▶ YouTube |
| Luce KVFlash ▶ YouTube |
The Lucebox Machine
Local AI should be the default, not a privilege. Private data, no per-token bill, no vendor lock-in. Lucebox pairs the R9700 with Strix Halo and ships this open engine ready to run.
See the hardware and current benchmarks at lucebox.com.
Request for Contributions
We welcome focused contributions to CUDA and HIP kernels, speculative inference, support for more consumer GPUs and APUs, performance benchmarks, and client harnesses.
Citation
@software{lucebox_2026, title = {Lucebox: Speculative inference for heterogeneous consumer hardware}, author = {Lucebox}, url = {https://github.com/Luce-Org/lucebox}, year = {2026} }
Community
- Discord: discord.gg/yHfswqZmJQ
- Website: lucebox.com
- Issues: github.com/Luce-Org/lucebox/issues
- Blog: lucebox.com/blog












