FLUX 3: One Multimodal Model | Black Forest Labs

5 min read Original article ↗

One multimodal model for ImageVideoAudioAction-Prediction
Creations are truer to life in every kind of style.

BurdaCanvaCloudflareComfyFalGensparkHiggsfieldKreaLovartMagnific Nous ResearchNvidiaOpenartOpenrouterPicsartRunwareRunwayEnvatoHeyGen

One model,
multiple modalities.

Stylistically diverse beyond just cinematic, with native audio and up to 20 second clips in a single generation. Start from text, an image, or keyframes, and get multiple shots in one take.

Introducing FLUX 3

FLUX 3 Video’s core capabilities.

Handles simple or complex prompts.

Draft mode.

Explore creative directions fast at a fraction of the cost. A draft generation returns a fast preview of your prompt. When a draft is right, send it back and FLUX 3 will render the same video at full quality.

Koi drifting between floating paper lanterns at night, draft render

Draft

Fast generation at a fraction of the cost.

Koi drifting between floating paper lanterns at night, full-quality render

Normal

Renders the draft in full quality.

Already in production.

What teams are saying about FLUX 3 Video.

  • FLUX 3 puts a new medium of entertainment and storytelling in our users’ hands, and it responds well to an agent’s direction. Hermes can generate a shot, checks it, then continues from it, so those 20-second clips become cohesive longform pieces.
    Nous ResearchDillon Rolnick, CEO, Nous Research
  • FLUX 3 is a significant addition to our next generation of video tools, including Picsart's flagship AI Playground. Having motion and audio grounded in the same model lets us treat each clip as a self-contained shot and handle narrative and timing at the pipeline level. Our users expect access to the latest and greatest models, and FLUX 3 raises the bar for what they can create with AI.
    PicsartMikayel Vardanyan, COO & Co-Founder, Picsart
  • With FLUX 3, we have the opportunity to boldly explore new directions and experiment with ideas that were previously unimaginable. The model delivers such high visual quality that we can bring our IPs to life in fully on-brand videos — without the time and cost constraints of traditional productions. This unlocks a whole new space for innovation and opens up entirely new possibilities.
    Burda MediaRebecca Gottwald, Chief AI Officer, Burda Media
  • Most video models are built for film. FLUX 3 goes beyond that, into the work our users also spend their days on: product, brand, motion design. That range in a single model is why we wanted it on Magnific from day one.
    MagnificOmar Pera, CPO, Magnific
  • We've been building with FLUX 3 ahead of launch, and it changes what our Creative Pros can do with video. The quality bar has moved to a point where the ideas that used to get cut for budget or time reasons are now on the table. We're seeing our community go from concept to polished, on-brand video in hours instead of weeks. For a platform built around empowering creatives, that's exactly the kind of step change we want to put in their hands.
    EnvatoHichame Assi, CEO, Envato

Use cases

Content

Content

Virtual-try-on

Virtual-try-on

Storyboarding

Storyboarding

cinematic western shot of a saddled horse galloping through the desert

Filmmaking

Advertising

Advertising

whiteboard explainer animation of a hand drawing dot clusters with a marker

Explainer Videos

street-food chef flame-cooking in a wok at a night market

Documentary

World simulation

World simulation

Worldbuilding

Worldbuilding

a wolf running through a snowy forest

Nature & Wildlife

anime-style schoolgirl on a rooftop overlooking a city at sunset

Anime

liquid droplets forming a logo reveal on black

VFX & Titles

a harpist performing inside a glowing blue ice cave

Music & Performance

retro CRT terminal booting with green monospace text

Text and Typography

Each modality makes the others better.

FLUX 3 builds on Self-Flow, our approach for efficiently aligning multimodal generation and understanding within the same underlying architecture.

Two-panel figure comparing Self-Flow with Flow Matching. Self-Flow comes out ahead in both: lower generation error across the modalities on the left, higher manipulation success rate through finetuning on the right.

Self-Flow vs. Flow Matching (FM). Left: generation error (Fréchet distance) per modality, each normalized to FM = 100 (lower is better). Right: success rate on manipulation tasks averaged over four task groups through finetuning (higher is better).

A growing collection of videos created by the FLUX 3 community.

Frequently asked questions

What is FLUX 3?

FLUX 3 is Black Forest Labs' multimodal foundation model for generating and understanding video, audio, images, and actions within one unified architecture.

What can I use as input?

FLUX 3 Video supports text-to-video, image-to-video, video-to-video, video continuation, and controlled transitions using keyframes.

How long and at what resolution can FLUX 3 generate video?

FLUX 3 can generate clips up to 20 seconds long in a single generation, in HD (up to 1 megapixel per frame) or FHD (up to 2 megapixels per frame). Resolution bands are set by total pixels per frame, not aspect ratio. Draft mode is HD only.

Does FLUX 3 generate video and audio together?

Yes. FLUX 3 can generate optional native audio with the video, including multilingual dialogue, synchronized speech, sound effects, and environmental ambience. Audio is included at no extra charge.

How much does FLUX 3 Video cost?

FLUX 3 uses pay-as-you-go pricing with no subscriptions or seat fees. Text or image to video costs $0.06/sec for Draft HD, $0.17/sec for standard HD, or $0.29/sec for standard FHD. Video to video costs $0.12/sec for Draft HD, $0.41/sec for standard HD, or $0.53/sec for standard FHD. Five seconds of standard text-to-video in HD costs $0.85. View pricing

What is FLUX 3 Draft mode?

Draft mode generates a faster, lower-cost HD preview. Once the direction is right, the draft can be enhanced into a full-quality FHD result. Draft Enhance is priced at the corresponding standard FHD rate.

How can I access FLUX 3?

FLUX 3 Video is available through the Black Forest Labs dashboard and API on a pay-as-you-go basis. Enterprise customers can request volume discounts, custom pricing, SLA guarantees, and dedicated support. Talk to sales

Are FLUX 3 Image, Action, and open weights available?

FLUX 3 Image, FLUX 3 Action, and the FLUX 3 Dev open-weight backbone are separate parts of the rollout. FLUX 3 Image generates and edits images, FLUX 3 Action is rolling out through selected research and commercial partners beginning with mimic robotics, and open-weight access to the FLUX 3 Dev multimodal backbone is coming soon. Read the rollout plan

Make things with FLUX 3.

One model built for exploration and production.