MiniMax H3 Super Acceleration | NVIDIA

7 min read Original article ↗

MiniMax H3 Super Acceleration fast draft generation and high-resolution refinement, powered by Sol Engine 6.85 s for a 5-second 768p video · 14.93 s for a 10-second video

H3 Super Acceleration first uses H3 with a LoRA to generate a four-step draft at 896×512. It then upsamples the draft and performs three LTX refinement steps at the target resolution with Sol-Attn. Combining the measured stages on one NVIDIA GB200 gives 22.2× speedup for a 5-second 1344×768 video and 27.7× speedup for a 10-second video over the published SGLang baseline.

* Equal contribution NVIDIA SANA Team

Production at a glance

Up to 378K videos every 30 days

One NVIDIA GB200 running H3 Super Acceleration continuously can turn the measured end-to-end latency into production-scale output.

768p · 5-second videos6.852 s measured E2E

12.6Kvideos / day17.5 hours of finished video

378Kvideos / 30 days525 hours of finished video

768p · 10-second videos14.931 s measured E2E

5.79Kvideos / day16.1 hours of finished video

174Kvideos / 30 days482 hours of finished video

Ideal full-utilization estimate. Counts assume one GB200, batch size 1, 24 × 7 serial inference, no idle time or failed jobs, and a 30-day month. Finished-video hours express media volume; encoded storage in GB or TB depends on the delivery codec and bitrate.

01 — Motivation

A practical speed–quality tradeoff for production video inference.

Our target is to reduce end-to-end inference time while keeping the visual and audio differences small enough for practical use. Instead of running many H3 steps at full resolution, H3 Super Acceleration performs most generation work in a short low-resolution draft and uses a lightweight high-resolution refinement pass.

This is not lossless acceleration. H3 Super Acceleration changes the sampling path, so it is not bit-exact with the SGLang baseline. Differences can appear in detail, texture, motion, or audio. The paired videos below let readers assess that tradeoff directly.


02 — Two-Stage Generation Pipeline

A low-resolution H3 draft followed by a short LTX refinement pass.

Each stage has one job. H3 creates the initial video at 896×512 in four denoising steps. LTX then upsamples and refines that draft at the requested output resolution in three steps with Sol-Attn. The two stages run serially on one GB200.

Hardware

1× GB200

Single-GPU inference

Stage 1 · 896×512

4 steps

H3 + LoRA · 24 FPS

Stage 2 · 768p / 1080p / 2K

3 steps

LTX · Sol-Attn · 24 FPS

Total

7 steps

Draft generation + refinement

Stage 1 · 896×512 · 24 FPS

Fast draft generation

4 steps

Prompt and conditioningGeneration inputs enter the H3 pipeline.

H3 + LoRAFour-step denoising at 896×512 uses the task-specific LoRA to generate the draft efficiently.

Draft videoComposition and motion are ready for refinement.

Stage 2 · 768p / 1080p / 2K · 24 FPS

High-resolution refinement

3 steps

Spatial upsamplingThe draft video is resized to the target resolution.

LTX refinementA three-step Sol-Attn pass restores high-resolution detail and consistency.

Final high-resolution videoThe refined result is decoded and returned as the final output.

H3 generates a 24 FPS draft at 896×512; after upsampling, LTX refines the same video at 24 FPS and 768p, 1080p, or 2K output resolution.

StageModelResolutionFPSDenoising stepsOutput
Stage 1MiniMax-H3 + LoRA896×512244Draft video
Stage 2LTX · Sol-Attn768p / 1080p / 2K243Final refined video

03 — Measured Latency

End-to-end 768p latency on one NVIDIA GB200.

The latest Sol-Super results use Sol-Attn for the complete Stage 2 service. Adding the measured Stage 1 and Stage 2 blocks gives 6.852 seconds for a 5-second 1344×768 video and 14.931 seconds for a 10-second video, corresponding to 22.2× and 27.7× speedup over the published SGLang baseline.

ResolutionDurationDiffusersSGLangSol EngineSol-Super
E2E
Sol-Super
vs. SGLang
1344×7685 s167.8 s152.3 s45.4 s6.852 s22.2×
1344×76810 s468.2 s414.1 s132.5 s14.931 s27.7×

Benchmark basis. Diffusers, SGLang, and Sol Engine are author-supplied warm end-to-end measurements. For the 5-second and 10-second settings, Sol-Super combines separately measured Stage 1 blocks of 4.1728 and 9.8204 seconds, respectively, with the latest Stage 2 means from 10 hot complete-service requests on one GB200; model loading and warmup are excluded. Sol-Super speedup is SGLang E2E ÷ Sol-Super E2E.

SGLang vs. Sol-Super end-to-end latency

1344×768 · 5 s

SGLangBaseline · measured

Sol-SuperStage 1 + Stage 2

S14.173 sS22.679 s

22.2× faster

1344×768 · 10 s

SGLangBaseline · measured

Sol-SuperStage 1 + Stage 2

S19.820 sS25.111 s

27.7× faster

The SGLang baseline fills the track in each setting. Sol-Super bars are enlarged only to keep the Stage 1 and Stage 2 labels readable; their internal proportions, printed latency values, and speedups are the measured results. Each red arrow ends at the right edge of its corresponding SGLang bar.

Our baseline vs. current Sol-Super

1344×768 · 5 s · one NVIDIA GB200

2.62× faster61.9% lower latency

Our baselineStage 1: 7.199 s · official H3 VAE decode · Dense eager Refiner · original LTX-2.5 Video VAE decode

H3 DiT · 3.772 s H3 VAE · 3.427 s Dense Refiner · 1.930 s LTX-2.5 VAE decode · 6.404 s

16.320 s

Gemma · 0.644 s LTX VAE encode · 0.143 s

Current Sol-SuperStage 1: 4.1728 s · TAEH3 decode · text/image encoding · Sol-Attn Refiner · LTX TAEHV final decode

TAEH3 decode · 0.028 s Stage 1 text/image encoding · 0.3728 s Gemma · 0.644 s LTX VAE encode · 0.143 s Sol-Attn Refiner · 1.079 s LTX TAEHV decode · 0.187 s

All segments use one linear scale. Labels are placed inside the bar wherever they fit; the compact keys below each bar list only the narrow segments. The red arrow spans from the current latency to the baseline endpoint. Totals exclude the shared mux and cleanup stage. In our baseline, Stage 1 combines 3.772 seconds of H3 DiT and shared work with 3.427 seconds of official H3 VAE decode, for 7.199 seconds total. Sol-Super Stage 1 combines the same 3.772 seconds of H3 work with 0.028 seconds of TAEH3 decode and 0.3728 seconds of text/image encoding, for 4.1728 seconds total. The decoder values are averages from two matched 1344×768 cases. Refiner blocks report complete Refiner latency; the Sol-Attn DiT portion takes 0.985 seconds. The original LTX-2.5 Video VAE decoder takes 6.404 seconds in the matched 121-frame decode benchmark.

Why the encoder stays unchanged. Stage 2 retains the original LTX-2.5 Video VAE encoder because the Refiner was trained on its latent distribution. LTX TAEHV is used only for the final decode, where it reduces latency without feeding shifted conditioning latents back into the Refiner.


04 — Visual Comparison

Paired visual comparisons plus standalone 1440p examples.

Each 768p and 1080p pair uses matching generation inputs to compare the SGLang baseline with H3 Super Acceleration. No matched 1440p baseline was recorded, so the last panel contains two standalone H3 Super Acceleration results and makes no speed claim.

Resolution 1344×768 · Duration 5 s

Scene — A man and woman exchange a tense remark on a crowded Japanese street.

SGLang BaselineReference result

H3 Super AccelerationAccelerated result

Resolution 1344×768 · Duration 10 s

Scene — A man speaks to a handheld camera on a quiet suburban street.

SGLang BaselineReference result

H3 Super AccelerationAccelerated result

Resolution 1080p · Duration 5 s

Scene — Two martial artists face one another in a bamboo forest.

SGLang BaselineReference result

H3 Super AccelerationAccelerated result

Resolution 1080p · Duration 10 s

Scene — A woman speaks to a handheld camera after a gym workout.

SGLang BaselineReference result

H3 Super AccelerationAccelerated result

Resolution 2560×1440 · Duration 10 s

Two standalone H3 Super Acceleration results; no matched baseline or speed claim is shown.

City Crowd ReactionH3 Super Acceleration

Futuristic WarriorH3 Super Acceleration


05 — Tokenomics

Throughput, revenue, and compute-adjusted margin for one fully utilized NVIDIA GB200.

MiniMax's official API pricing lists 768P H3 output at $0.08 per second: $0.40 for a 5-second video and $0.80 for a 10-second video. Under the clearly stated assumptions below, the combined Sol-Super latency corresponds to a 97.1%–97.4% GPU-only gross margin.

Setting / priceSGLang
videos / hour
Sol-Super
videos / hour
SGLang
revenue / hour
Sol-Super
revenue / hour
Speedup
vs. SGLang
Sol-Super
gross margin
768p · 5 s$0.40 / video23.64525.41$9.46$210.1622.2×97.4%
768p · 10 s$0.80 / video8.69241.10$6.95$192.8827.7×97.1%

Ideal full-utilization scenario. We assume one GB200 runs continuously at 100% utilization and costs $5.50 per GPU-hour under a one-year Long-Term Agreement. Videos per hour are 3,600 ÷ measured E2E latency; hourly revenue is throughput × price per video; gross margin is (revenue − $5.50) ÷ revenue. This is a GPU-only margin: it excludes idle time, failed jobs, storage, networking, staffing, software, and other operating costs, so it is not net profit.