Speed-of-Light MiniMax-H3 on Single NVIDIA DGX Spark Blackwell Superchip

9 min read Original article ↗

Sol-Engine × Sol-Attn · MiniMax-H3 Inference

Sol-H3- Spark Accelerating MiniMax-H3 768p Video Generation on a Single NVIDIA DGX Spark in 1 Minute

NVIDIA Research, Efficient AI Team & Singapore Lab.

768p resolution 5s duration 56s e2e latency

A two-stage pipeline specialized for a single NVIDIA DGX Spark generates a 384p draft and refines it to 768p.

Sol-H3 code - Spark Sol-Engine page

Sol-H3 introduction 1344×768 / 24 FPS

Sol-H3 on Spark · Introduction · 1344 × 768 · 24 FPS

01 / Results

Less than one minute

56 seconds from a fresh prompt to a playable 768p video with T2VA or FL2VA. Ref2VA is slightly slower with more tokens.

Complete request latency

Timing notes & breakdown

The two hot measurements reflect production serving: from prompt arrival, including fresh Qwen encoding, to a completed MP4 with audio.

Where the 56.17 s goes

Share of E2E latency

35.2% 44.0% 12.5%

  1. Qwen prompt encoding2.4%
  2. H3 4-step generation35.2%
  3. H3 upscaler + VAE Adapter3.8%
  4. LTX 3-step refinement44.0%
  5. LTX VAE decode12.5%
  6. MP4 output + audio mux1.3%
  7. Request overhead0.9%

02 / Generations

Sol-H3 showcase

Selected generations across T2VA, FL2VA and Ref2VA.

Example 01 · Stop-motion animation

View exact prompt

subject_definitions: <Subject 1> is the original grey felt raccoon in <Picture 1>, with the same wool face, dark eye patches, small paws and teal apron. <Picture 1> also defines the handmade cream-and-pistachio miniature pastry shop, counter and tart topped with one red strawberry. Preserve these character, material, prop and environment references in a new animated moment. detailed_description: Core concept: a tiny thief's innocent expression is betrayed by the evidence still in its paw. Create a five-second tactile stop-motion scene in a wide landscape composition. Show visible felt fibers, soft miniature lighting and deliberately stepped yet coherent puppet movement. Shot description: one continuous, locked medium-wide view containing the raccoon, tart and counter. From 0.0 to 1.5 seconds, <Subject 1> leans slightly toward the tart, glances sideways and slowly extends one paw toward the single strawberry. From 1.5 to 3.2 seconds, the paw gently closes around the berry and lifts it clear of the tart in one readable motion. Keep the strawberry's red body and green leaves visible, and leave the tart intact on the counter. From 3.2 to 5.0 seconds, a small offscreen wooden creak makes the raccoon straighten and look directly toward the camera with exaggerated innocence. It briefly holds that pose while one ear twitches and the apron settles; the stolen berry remains plainly visible in its raised paw. Maintain one raccoon and one strawberry throughout. No cuts, floating props, readable text, logos, subtitles, reference-image flashes or freeze frames. overall_soundscape: Tiny felt-and-fabric rustles, a faint countertop tap and one quiet wooden creak followed by soft shop room tone. No speech. non_diegetic_music: None.

Example 02 · Mandarin dialogue

View exact prompt

subject_definitions: <Subject 1> is the Chinese mother in her early sixties in <Picture 1>, preserving her kind face, short salt-and-pepper hair, sage-green cardigan and ivory blouse. <Picture 1> also supplies the single small walnut tabletop radio with cream speaker fabric and two knobs, the oak table and warm quiet home. <Subject 2> is her adult daughter around thirty in <Picture 2>, preserving her face, straight dark shoulder-length hair, terracotta overshirt and off-white top. Compose these two women sitting together at that one table. Use the pictures as separate identity references, not as cutaway shots. detailed_description: Core concept: a repaired old radio creates a warm exchange between a mother and her adult daughter. A five-second photorealistic film, one continuous steady eye-level medium two-shot, both women's faces and the single radio visible, soft afternoon window light, clean composition and natural skin texture. In the first two seconds, the daughter looks at her mother with curiosity and clearly asks in standard Mandarin Chinese, “妈,修好了?” The mother looks up from the radio and answers warmly in standard Mandarin Chinese, “好了,听听看。” Her hand lightly rests by a knob; after finishing her reply, she turns it once with a small natural finger motion. They exchange a small smile. Maintain each speaker's identity and clothing, one radio on the table, no other people or distracting props. Each woman moves her lips only during her own line. Complete both short lines before the final moment, without overlap. No cuts, subtitles, text, logos, reference-picture flashes or freeze frames. overall_soundscape: Clear foreground Mandarin dialogue with two distinct female voices. Daughter: “妈,修好了?” Mother: “好了,听听看。” Speak only these Chinese words, with natural Mandarin pronunciation and no English, no narrator or extra chatter. Keep room ambience very quiet beneath speech. After the reply, a soft radio-knob click and a faint radio hiss; no broadcast speech or song. non_diegetic_music: None.

03 / Approach

Coarse-to-fine pipeline

Generate the scene on a compact H3 latent canvas, upscale in H3 space, then transfer to LTX for three-step refinement and full Conv VAE decoding.

01 · Generate

MiniMax H3 · 384p

672 × 384 · 124 frames

NVFP4 Qwen · FP8 DiT · 4 steps LoRA

Latent transfer

VAE Adapter

H3 space → LTX space

VAE-free latent mapping

02 · Refine & decode

LTX · 768p

1344 × 768 · 121 frames

LTX-2.5 · Sol-Attn · 3 steps · no text encoder

Stage 1 supports compatible few-step LoRAs and 384p, 480p or 512p drafts, with matching latent upscaling to 768p.

04 / Memory

The “black magic”: How Sol-H3 gets faster

Two models need more than faster kernels: they need room to stay loaded. Our estimated naive two-stage configuration exceeds Spark’s shared memory budget.
Two changes make residency practical: (1) a learned VAE Adapter removes inter-stage video decode/re-encode; (2) a cached generic refinement prompt removes online Gemma and connector processing. Quantization then lets the retained components stay resident. The latent carries the scene condition; the generic prompt describes refinement quality, not scene content.

Resident memory · GiB

Memory breakdown

DGX Spark exposes CPU and GPU unified memory as one 119.68 GiB pool. Crossing the line means the complete pipeline cannot stay resident.

The full official MiniMax-H3 pipeline cannot fit entirely in DGX Spark’s unified memory. A naive two-stage pipeline requires both complete model stacks, exceeding the memory budget even with low-bit weight quantization. Sol-H3 removes redundant components through VAE-free latent transfer and refiner prompt caching, allowing the two-stage pipeline to stay fully resident and run without out-of-memory errors.

Technical details 5 optimizations

Sol-Attn

Sparsify on the fly

Query-dependent · one-pass · training-free

Select blocks while attention runs.

Mechanism

Each query block sets its own threshold over lightweight query–key scores. Selected blocks use exact attention; skipped blocks receive an approximate correction from pooled K/V summaries.

Why it is faster

Routing, sparse attention and approximate correction share one online-softmax pass. Sol-Attn writes no separate full score map or routing-index tensor and needs no retraining.

VAE Adapter

Direct latent transfer

Direct H3 → LTX mapping VAE-free transfer

Skip the pixel-space handoff.

Mechanism

A learned VAE Adapter maps the upscaled H3 latent directly into LTX space. The pipeline never decodes the Stage 1 latent to frames and re-encodes it for Stage 2.

Why it is faster

Removing the H3 VAE decoder and LTX VAE encoder eliminates two large resident components and both inter-stage forwards. Only the final LTX VAE decoder remains.

Refiner conditioning

Cache the refiner prompt

Generic prompt · encoded once

Let the latent carry the scene.

Mechanism

The 384p draft latent already gives the refiner a strong content and motion prior. A content-agnostic quality prompt can be encoded once through Gemma and the connector.

Why it is faster

The cached post-connector features remove runtime LTX text encoding and let Gemma stay out of the resident pipeline. Stage 1 still runs a fresh Qwen encode for every request.

Resident deployment

Keep the runtime resident

Quantize · fit · stay warm

Fit both stages into one shared pool.

Mechanism

NVFP4 Qwen, an FP8 H3 DiT, the VAE Adapter, cached Stage 2 context and low-bit video components reduce the resident footprint while preserving the fixed two-stage recipe.

Why it is faster

A resident pipeline avoids per-request model construction, checkpoint parsing and weight switching. Each request moves directly from prompt encoding through both video stages.

Kernel & I/O fusion

Fuse the execution path

BSA · layout · encode

Remove redundant work end to end.

Mechanism

Stage 1 VSA runs selected blocks through cuDNN BSA with fused routing and merge. The final Conv VAE emits its preferred NHWC layout directly, while chunked encode and audio mux avoid extra staging.

Why it is faster

Eliminating dense fallbacks, standalone transposes and unnecessary host synchronization reduces memory traffic and keeps delivery work inside the measured request path.

Credits

Acknowledgements

We thank the teams and open projects that helped turn Sol-H3 into one end-to-end Spark pipeline.

Core contributors Haopeng Li Junsong Chen Yitong Li Jincheng Yu Jingyu Xin Haocheng Xi Song Han Enze Xie
MiniMaxMiniMax-H3 base weights, used with FP8 quantization in Stage 1, and native video-and-audio generation.
Sol-Engine & Sol-AttnThe inference framework, SuperH3 workflow and sparse refinement attention.
Comfy-OrgPre-quantized MiniMax-H3 FP8 checkpoints used in our earlier Spark experiments.
LightricksLTX-2.5 refinement and the official convolutional video VAE.
LBH-123-AIThe learned H3 latent upscaler used before the VAE adapter.
LightX2VOpen H3 acceleration work supporting the original Sol-Engine workflow.
Video DeltaNet · UC BerkeleyComplementary hybrid-attention work for MiniMax-H3.
Humanize + KDA Operator and kernel optimization alongside Sol-Engine.
FastH3 · Hao AI Lab @ UCSDThe four-step FastH3 VSA model used for our 384p draft stage.

Citations

If this project helps your work, please cite the Spark release and the underlying Sol-Engine and Sol-Attn papers.

BibTeX · Release + core papers

@misc{solh3spark2026,
  title        = {Sol-H3: Speed-of-Light MiniMax-H3 on Single NVIDIA DGX Spark Blackwell Superchip},
  author       = {{Sol-H3 Team}},
  year         = {2026},
  howpublished = {\url{https://nvlabs.github.io/Sana/Sol-Engine/Sol-H3-Spark/}},
  note         = {Single-Spark project release}
}
@misc{li2026solvideoinferenceengine,
  title         = {Sol Video Inference Engine: Agent-Native Full-Stack Acceleration Framework for Efficient Video Generation},
  author        = {Yitong Li and Junsong Chen and Haopeng Li and Haozhe Liu and Jincheng Yu and Ligeng Zhu and Ping Luo and Song Han and Enze Xie},
  year          = {2026},
  eprint        = {2606.23743},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CV},
  doi           = {10.48550/arXiv.2606.23743},
  url           = {https://arxiv.org/abs/2606.23743}
}
@misc{li2026solattn,
  title         = {Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification},
  author        = {Haopeng Li and Yitong Li and Junsong Chen and Tian Ye and Haozhe Liu and Jincheng Yu and Duomin Wang and Ruihua Zhang and Zeke Xie and Enze Xie and Song Han},
  year          = {2026},
  eprint        = {2607.24027},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CV},
  doi           = {10.48550/arXiv.2607.24027},
  url           = {https://arxiv.org/abs/2607.24027}
}