AERIOXFLUX
AI Tools
AI Tools · image video models

FLUX 3 Collapses Image, Video, Audio and Robot Actions Into One Model

Black Forest Labs shipped a single set of weights that generates 20-second video with native audio, edits images, and predicts robot manipulation — and the last item is the one that matters.

Flux Desk·2026-07-26·5 min read

Black Forest Labs released FLUX 3 on July 23, and the specification reads like three separate product launches stapled together: a video model that generates up to 20 seconds of clip with native audio, an image model that synthesizes and edits across styles and resolutions, and an action model that predicts robot manipulation. The temptation is to grade each one against its category leader. That misses what the company actually shipped. FLUX 3 is one multimodal foundation model — images, video, and audio learned jointly, on a single flow-matching architecture the lab calls Self-Flow — and the three products are surfaces on the same weights.

That distinction is the whole story.

Native audio is not a feature, it's evidence

Most video generators that produce sound produce it in a second pass: generate pixels, then hand the pixels to an audio model that guesses what they should sound like. The result is competent and slightly wrong in a way that is hard to name — footsteps that land a beat off, an impact whose sound has the right timbre and the wrong weight.

FLUX 3 generates audio and video from the same model, and Black Forest Labs specifically claims strength in associating sound with physical events. That claim is only interesting because of how it must have been earned. To place the right sound at the right frame, a model has to have learned something about the event underneath both — that this is a door closing, with this much mass, on this surface. Audio-visual synchrony is not a rendering trick. It is a side effect of having a representation that outlives any single modality.

The company also cites facial expressions, character consistency, and multilingual dialogue as strengths, which points at the same underlying thing: identity and intent held stable across time rather than re-derived frame by frame.

The numbers, and what they cover

Black Forest Labs published preliminary head-to-head preference results — human raters, 720p, 10-second clips. FLUX 3 was preferred over Luma Ray 3.2 (93%), Runway Gen-4.5 (77%), Grok Imagine Video (69%), and Kling v3 Pro (60%).

Read the spread rather than the headline. A 93% preference rate is a category gap; 60% against Kling v3 Pro is a real but ordinary win, the kind that reverses with a prompt distribution change or a checkpoint update. The lab is describing itself as clearly ahead of the middle of the field and modestly ahead of the strongest competitor — which is roughly what an honest evaluation looks like, and notably not what a marketing page usually says.

Two caveats belong in the same breath. These are the lab's own preliminary evaluations, not third-party benchmarks. And they were run at 10 seconds and 720p, while the headline capability is 20 seconds — with chaining to reach sequences of several minutes. The clips people will actually judge FLUX 3 on are longer than the clips it was measured on.

FLUX 3 Action, and the reason this launch is different

The part of the announcement least likely to trend is the part most likely to matter. FLUX 3 integrates action prediction natively, and Black Forest Labs is developing FLUX-mimic with mimic robotics for dexterous manipulation. The video backbone — the thing trained to predict what happens next in a scene — is being adapted to predict what a robot should do next.

This is the convergence the robotics field has been circling for two years. A generative video model that can produce twenty coherent seconds of a hand picking up a cup has, in some functional sense, learned how a hand picks up a cup: the approach, the contact, the load transfer, what the scene looks like after. Control policies need exactly that. The bet is that the expensive part of robot learning — a usable model of physical consequence — can be paid for once with internet-scale video and then specialized, rather than collected teleoperation episode by teleoperation episode.

Nvidia has been arguing this with Cosmos. Google DeepMind has been arguing it with Genie. What is different here is the packaging: not a research world model with a robotics story attached, but a commercial media generator whose robotics variant ships through an industrial partner. When the same weights that render an ad are also steering a gripper, the economics of physical AI start looking less like a research program and more like a product line with two revenue streams.

The catch: almost none of this is available

FLUX 3 Video and FLUX 3 Action are in early access, gated behind a request form. Image access is promised "in the coming weeks." APIs, model weights, and the open-weight FLUX 3 Dev release are staged over the coming months. Action reaches developers through partners, not directly.

Black Forest Labs earned its reputation on open weights — FLUX.1 Dev became the default fine-tuning substrate for an entire generation of image tooling because anyone could take it and run. FLUX 3 inverts that sequence: capability demonstrated first, access rationed, weights last. It is the same pattern every lab adopts as its models get expensive enough to be worth protecting, and it is the reason a launch post is a claim rather than a fact until the queue clears.

What to actually watch

Three things will settle whether this launch was a step change or a strong quarter.

Whether 20 seconds holds up. The evaluations were run at 10. Coherence degradation is superlinear, and the second half of a long generation is where models reveal what they never learned.

Whether chaining is seamless. "Sequences of several minutes" via chaining is either a genuine long-form capability or a stitching workflow with visible seams. That difference decides whether this is a tool for clips or a tool for films.

Whether FLUX-mimic produces a real deployment. An industrial partnership on a robotics variant is the strongest signal in the announcement, and also the easiest to quietly not mention again.

The unified-architecture claim is the one worth taking seriously regardless. Every lab now sells a bundle of modalities; very few sell one model that learned them together. If the audio-physical grounding and the action transfer both hold, the interesting question stops being who makes the best video model and becomes who owns the general-purpose predictor that video, sound, and motion all fall out of.

#black-forest-labs#flux-3#video-generation#world-models#robotics

The state of AI, in flux.

The directory + magazine for AI tools and the workflows people use to make money with them.

🔥 The Sauce Drop

The week's highest-earning AI workflows, in your inbox.

Some outbound links are affiliate links — Flux may earn a commission at no cost to you; this never affects rankings. Earnings figures are self-reported and not guarantees of income; most people earn less, some earn nothing.