Generative AI

FLUX 3: One Backbone for Video, Images, Audio — and Robots

Black Forest Labs launched FLUX 3, a multimodal model with 20-second native-audio video, plus FLUX-mimic, a robot action model built on the same backbone.

FLUX 3: One Backbone for Video, Images, Audio — and Robots — article cover
On this page6 SECTIONS
  1. One Backbone, Three Modalities: The Self-Flow Bet
  2. Video Capabilities and Early Benchmarks
  3. FLUX-mimic: A Model That Understands Physics, Controlling Robots
  4. Availability: Early Access, APIs, and an Open-Weight Plan
  5. What It Means for Creators and Engineering Teams
  6. Sources

Black Forest Labs announced FLUX 3 on July 23, 2026: its first multimodal foundation model, in Early Access now, and the first built entirely on the company’s “Self-Flow” approach. The core claim: video, images, and audio are each “a projection of the same underlying reality,” so instead of training three separate tasks, one architecture should learn generation and understanding jointly, treating every modality as evidence about a single world.

The second post published the same day may matter more in the long run. FLUX-mimic, a video-action model built on the FLUX 3 backbone, is already doing real work on an Audi production line. A generative-media company crossing into robot control is a bigger signal than any benchmark score.

One Backbone, Three Modalities: The Self-Flow Bet

FLUX 3 was trained on tens of millions of hours of general video plus manipulation-focused footage, with video prediction consuming over 95% of the training compute. Self-Flow aligns multimodal generation and understanding in a single architecture, scaled up with significantly more compute and data than the company’s earlier image-only line. The wager is that understanding and generation are two uses of one world model, not two products. One detail shows what the bet costs: when action prediction was added to training, video quality dropped by up to 10% — and fully recovered after 3,500 training steps. The same backbone can carry both content creation and robot control.

Video Capabilities and Early Benchmarks

Video is the headline: clips up to 20 seconds with native audio, covering text-to-video, image-to-video (animation or reference-based), video-to-video with character carryover, keyframe transitions, multilingual dialogue, agentic multi-shot chaining, and typography strong enough for animated design work. Image synthesis and editing round out the launch, with better handling of complex prompts and multilingual text rendering.

The benchmarks come with a preliminary label, correctly. Human raters compared 10-second 720p text-to-video clips with audio: FLUX 3 was preferred over Runway Gen-4.5 in 77% of comparisons, over Luma Ray 3.2 in 93%, up to 69% against Grok Imagine Video, 60% against Kling v3 Pro, and 52% against both Seedance 2.0 and Gemini Omni Flash.

FLUX-mimic: A Model That Understands Physics, Controlling Robots

FLUX-mimic is a collaboration with the robotics company mimic robotics: a lightweight action decoder attached to the FLUX 3 backbone reads robot actions directly out of the same world representation the model uses to generate video. The logic is that a model capable of generating realistic video must already understand contact, motion, weight, and cause and effect — and that understanding can be repurposed for control.

The numbers: the backbone runs under 80 ms on a single RTX 5090, for a full system the authors say reacts in 101 ms. It beats prior vision-language-action models even with a frozen backbone, needs far less demonstration data, and recovers from failed grasps without being shown how. Audi has tested it on production tasks — kitting parts into trays, assembling electronic control units, and handling soft materials like seals and cables — work the company says conventional robotics could not do. Christoph Schneider of the Audi Production Lab credited the robots with solving “complex soft-body manipulation work that would have been simply impossible with conventional robotics.”

Availability: Early Access, APIs, and an Open-Weight Plan

FLUX 3 Video is in Early Access now; image Early Access is expected in the following weeks. The launch plan covers APIs and private weights for video and image, partner access for action starting with mimic robotics, and an open-weight “FLUX 3 Dev” backbone. For developers used to running weights on their own hardware, that last item is the commitment that matters most.

What It Means for Creators and Engineering Teams

For creators, the practical shift is audio-native video with usable typography and multi-shot control, which cuts dubbing and storyboarding work; dialogue in multiple languages generated with the clip removes an entire localization pass. For engineering teams, Self-Flow and FLUX-mimic point at a path: instead of maintaining separate models for video generation and robot control, one visually grounded backbone may serve both. When the open-weight release lands, and under what license, will decide how fast this ecosystem grows.

Sources

AI-assisted summary compiled from the sources above, reviewed by a human before publishing.

SHAREXEMAIL