Sol-Engine × Sol-Attn · MiniMax-H3 Inference
Sol-H3- Spark Accelerating MiniMax-H3 768p Video Generation on a Single NVIDIA DGX Spark in 1 Minute
NVIDIA Research, Efficient AI Team & Singapore Lab.
768p resolution 5s duration 56s e2e latency
A two-stage pipeline specialized for a single NVIDIA DGX Spark generates a 384p draft and refines it to 768p.
Sol-H3 introduction 1344×768 / 24 FPS
Sol-H3 on Spark · Introduction · 1344 × 768 · 24 FPS
01 / Results
Less than one minute
56 seconds from a fresh prompt to a playable 768p video with T2VA or FL2VA. Ref2VA is slightly slower with more tokens.
Timing notes & breakdown
The two hot measurements reflect production serving: from prompt arrival, including fresh Qwen encoding, to a completed MP4 with audio.
Where the 56.17 s goes
Share of E2E latency35.2% 44.0% 12.5%
- Qwen prompt encoding2.4%
- H3 4-step generation35.2%
- H3 upscaler + VAE Adapter3.8%
- LTX 3-step refinement44.0%
- LTX VAE decode12.5%
- MP4 output + audio mux1.3%
- Request overhead0.9%
02 / Generations
Sol-H3 showcase
Selected generations across T2VA, FL2VA and Ref2VA.
Example 01 · Stop-motion animation
View exact prompt
subject_definitions: <Subject 1> is the original grey felt raccoon in <Picture 1>, with the same wool face, dark eye patches, small paws and teal apron. <Picture 1> also defines the handmade cream-and-pistachio miniature pastry shop, counter and tart topped with one red strawberry. Preserve these character, material, prop and environment references in a new animated moment. detailed_description: Core concept: a tiny thief's innocent expression is betrayed by the evidence still in its paw. Create a five-second tactile stop-motion scene in a wide landscape composition. Show visible felt fibers, soft miniature lighting and deliberately stepped yet coherent puppet movement. Shot description: one continuous, locked medium-wide view containing the raccoon, tart and counter. From 0.0 to 1.5 seconds, <Subject 1> leans slightly toward the tart, glances sideways and slowly extends one paw toward the single strawberry. From 1.5 to 3.2 seconds, the paw gently closes around the berry and lifts it clear of the tart in one readable motion. Keep the strawberry's red body and green leaves visible, and leave the tart intact on the counter. From 3.2 to 5.0 seconds, a small offscreen wooden creak makes the raccoon straighten and look directly toward the camera with exaggerated innocence. It briefly holds that pose while one ear twitches and the apron settles; the stolen berry remains plainly visible in its raised paw. Maintain one raccoon and one strawberry throughout. No cuts, floating props, readable text, logos, subtitles, reference-image flashes or freeze frames. overall_soundscape: Tiny felt-and-fabric rustles, a faint countertop tap and one quiet wooden creak followed by soft shop room tone. No speech. non_diegetic_music: None.
Example 02 · Mandarin dialogue
View exact prompt
subject_definitions: <Subject 1> is the Chinese mother in her early sixties in <Picture 1>, preserving her kind face, short salt-and-pepper hair, sage-green cardigan and ivory blouse. <Picture 1> also supplies the single small walnut tabletop radio with cream speaker fabric and two knobs, the oak table and warm quiet home. <Subject 2> is her adult daughter around thirty in <Picture 2>, preserving her face, straight dark shoulder-length hair, terracotta overshirt and off-white top. Compose these two women sitting together at that one table. Use the pictures as separate identity references, not as cutaway shots. detailed_description: Core concept: a repaired old radio creates a warm exchange between a mother and her adult daughter. A five-second photorealistic film, one continuous steady eye-level medium two-shot, both women's faces and the single radio visible, soft afternoon window light, clean composition and natural skin texture. In the first two seconds, the daughter looks at her mother with curiosity and clearly asks in standard Mandarin Chinese, “妈,修好了?” The mother looks up from the radio and answers warmly in standard Mandarin Chinese, “好了,听听看。” Her hand lightly rests by a knob; after finishing her reply, she turns it once with a small natural finger motion. They exchange a small smile. Maintain each speaker's identity and clothing, one radio on the table, no other people or distracting props. Each woman moves her lips only during her own line. Complete both short lines before the final moment, without overlap. No cuts, subtitles, text, logos, reference-picture flashes or freeze frames. overall_soundscape: Clear foreground Mandarin dialogue with two distinct female voices. Daughter: “妈,修好了?” Mother: “好了,听听看。” Speak only these Chinese words, with natural Mandarin pronunciation and no English, no narrator or extra chatter. Keep room ambience very quiet beneath speech. After the reply, a soft radio-knob click and a faint radio hiss; no broadcast speech or song. non_diegetic_music: None.
03 / Approach
Coarse-to-fine pipeline
Generate the scene on a compact H3 latent canvas, upscale in H3 space, then transfer to LTX for three-step refinement and full Conv VAE decoding.
MiniMax H3 · 384p
672 × 384 · 124 frames
NVFP4 Qwen · FP8 DiT · 4 steps LoRA
VAE Adapter
H3 space → LTX space
VAE-free latent mapping
LTX · 768p
1344 × 768 · 121 frames
LTX-2.5 · Sol-Attn · 3 steps · no text encoder
Stage 1 supports compatible few-step LoRAs and 384p, 480p or 512p drafts, with matching latent upscaling to 768p.
04 / Memory
The “black magic”: How Sol-H3 gets faster
Two models need more than faster kernels: they need room to stay loaded. Our estimated naive two-stage configuration exceeds Spark’s shared memory budget.
Two changes make residency practical: (1) a learned VAE Adapter removes inter-stage video decode/re-encode; (2) a cached generic refinement prompt removes online Gemma and connector processing. Quantization then lets the retained components stay resident. The latent carries the scene condition; the generic prompt describes refinement quality, not scene content.
Resident memory · GiB
Memory breakdown
DGX Spark exposes CPU and GPU unified memory as one 119.68 GiB pool. Crossing the line means the complete pipeline cannot stay resident.
The full official MiniMax-H3 pipeline cannot fit entirely in DGX Spark’s unified memory. A naive two-stage pipeline requires both complete model stacks, exceeding the memory budget even with low-bit weight quantization. Sol-H3 removes redundant components through VAE-free latent transfer and refiner prompt caching, allowing the two-stage pipeline to stay fully resident and run without out-of-memory errors.
Technical details 5 optimizations
Sparsify on the fly
Query-dependent · one-pass · training-free
Select blocks while attention runs.
Mechanism
Each query block sets its own threshold over lightweight query–key scores. Selected blocks use exact attention; skipped blocks receive an approximate correction from pooled K/V summaries.
Why it is faster
Routing, sparse attention and approximate correction share one online-softmax pass. Sol-Attn writes no separate full score map or routing-index tensor and needs no retraining.
Direct latent transfer
Direct H3 → LTX mapping VAE-free transfer
Skip the pixel-space handoff.
Mechanism
A learned VAE Adapter maps the upscaled H3 latent directly into LTX space. The pipeline never decodes the Stage 1 latent to frames and re-encodes it for Stage 2.
Why it is faster
Removing the H3 VAE decoder and LTX VAE encoder eliminates two large resident components and both inter-stage forwards. Only the final LTX VAE decoder remains.
Cache the refiner prompt
Generic prompt · encoded once
Let the latent carry the scene.
Mechanism
The 384p draft latent already gives the refiner a strong content and motion prior. A content-agnostic quality prompt can be encoded once through Gemma and the connector.
Why it is faster
The cached post-connector features remove runtime LTX text encoding and let Gemma stay out of the resident pipeline. Stage 1 still runs a fresh Qwen encode for every request.
Keep the runtime resident
Quantize · fit · stay warm
Fit both stages into one shared pool.
Mechanism
NVFP4 Qwen, an FP8 H3 DiT, the VAE Adapter, cached Stage 2 context and low-bit video components reduce the resident footprint while preserving the fixed two-stage recipe.
Why it is faster
A resident pipeline avoids per-request model construction, checkpoint parsing and weight switching. Each request moves directly from prompt encoding through both video stages.
Fuse the execution path
BSA · layout · encode
Remove redundant work end to end.
Mechanism
Stage 1 VSA runs selected blocks through cuDNN BSA with fused routing and merge. The final Conv VAE emits its preferred NHWC layout directly, while chunked encode and audio mux avoid extra staging.
Why it is faster
Eliminating dense fallbacks, standalone transposes and unnecessary host synchronization reduces memory traffic and keeps delivery work inside the measured request path.
Acknowledgements
We thank the teams and open projects that helped turn Sol-H3 into one end-to-end Spark pipeline.
| Core contributors | Haopeng Li Junsong Chen Yitong Li Jincheng Yu Jingyu Xin Haocheng Xi Song Han Enze Xie |
|---|---|
| MiniMax | MiniMax-H3 base weights, used with FP8 quantization in Stage 1, and native video-and-audio generation. |
| Sol-Engine & Sol-Attn | The inference framework, SuperH3 workflow and sparse refinement attention. |
| Comfy-Org | Pre-quantized MiniMax-H3 FP8 checkpoints used in our earlier Spark experiments. |
| Lightricks | LTX-2.5 refinement and the official convolutional video VAE. |
| LBH-123-AI | The learned H3 latent upscaler used before the VAE adapter. |
| LightX2V | Open H3 acceleration work supporting the original Sol-Engine workflow. |
| Video DeltaNet · UC Berkeley | Complementary hybrid-attention work for MiniMax-H3. |
| Humanize + KDA | Operator and kernel optimization alongside Sol-Engine. |
| FastH3 · Hao AI Lab @ UCSD | The four-step FastH3 VSA model used for our 384p draft stage. |
Citations
If this project helps your work, please cite the Spark release and the underlying Sol-Engine and Sol-Attn papers.
BibTeX · Release + core papers
@misc{solh3spark2026,
title = {Sol-H3: Speed-of-Light MiniMax-H3 on Single NVIDIA DGX Spark Blackwell Superchip},
author = {{Sol-H3 Team}},
year = {2026},
howpublished = {\url{https://nvlabs.github.io/Sana/Sol-Engine/Sol-H3-Spark/}},
note = {Single-Spark project release}
}
@misc{li2026solvideoinferenceengine,
title = {Sol Video Inference Engine: Agent-Native Full-Stack Acceleration Framework for Efficient Video Generation},
author = {Yitong Li and Junsong Chen and Haopeng Li and Haozhe Liu and Jincheng Yu and Ligeng Zhu and Ping Luo and Song Han and Enze Xie},
year = {2026},
eprint = {2606.23743},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
doi = {10.48550/arXiv.2606.23743},
url = {https://arxiv.org/abs/2606.23743}
}
@misc{li2026solattn,
title = {Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification},
author = {Haopeng Li and Yitong Li and Junsong Chen and Tian Ye and Haozhe Liu and Jincheng Yu and Duomin Wang and Ruihua Zhang and Zeke Xie and Enze Xie and Song Han},
year = {2026},
eprint = {2607.24027},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
doi = {10.48550/arXiv.2607.24027},
url = {https://arxiv.org/abs/2607.24027}
}