←Back to News
RESEARCH—Jul 23, 2026—10 Min Read

One Model, Four Modalities, and a Waiting Room

Model FLUX 3 (Video · Image · Action · Dev) · Launched 23 July 2026 · Team Pindow · 10 min

Pindow and FLUX 3

FLUX 3 is the most architecturally ambitious release of the quarter. Two months on, half of it still isn’t available, and the half that is has changed how we think about generated audio.

On 23 July, Black Forest Labs announced FLUX 3 from Freiburg: a single unified architecture trained jointly across image, video, audio and robot action prediction.

Read that list again, because it is genuinely strange. Three of those are things a filmmaker uses. The fourth is a robot deciding how to pick up a piece of soft material on a car production line.

The instinct is to treat the robotics tier as a distraction from the creative tiers, a research flourish bolted on for the investors. We think that reading is wrong, and the reason why is the most interesting idea in this entire series of posts. So let us take it seriously.

What actually shipped

Announced 23 July 2026. Built on BFL’s “Self-Flow” approach, meaning multimodal flow matching in a unified architecture that learns across modalities jointly rather than chaining specialist models together.

FLUX 3 Video. Gated early access at announcement, with a public API following in early August.

  • Up to 20 seconds per generation.
  • Native audio, with synchronised dialogue, sound effects and ambient generated in the same pass.
  • Multilingual dialogue. Generated in a language you choose.
  • Input modes covering text-to-video, image-to-video, video-to-video and keyframe-to-video.
  • Multi-shot handled by agentic chaining rather than a single extended pass. Worth knowing, because it is a different mechanism from native multi-shot modelling and it behaves differently.
  • Benchmarks were run at 720p. Full resolution capability was not disclosed at announcement.

FLUX 3 Image. Announced as coming in the weeks after launch. At the time of writing there is still no public image endpoint and no model ID, with access remaining behind an application-gated early-access form. We are noting this without editorialising further. It is simply the state of play, and anyone planning a pipeline around it should plan around its absence.

FLUX 3 Action. Via selected partners, starting with a Zurich robotics firm. The joint video-action model is reportedly in testing with a European car manufacturer for soft-body manufacturing tasks, with a reaction time around 101 milliseconds.

FLUX 3 Dev. An open-weight multimodal backbone spanning video, audio, image and action, explicitly the last tier to ship, targeted later in 2026. Licence terms undetermined.

Benchmarks, vendor-measured and labelled as such. BFL’s own preliminary evaluation on an early candidate, tested on 10-second 720p clips, reported FLUX 3 Video preferred in up to 77 percent of comparisons against competing models, with a published breakdown running from 93 percent against one rival down to an effective tie against two others. BFL themselves describe the results as preliminary. No independent verification was available at the time of writing.

Why the robot matters

Here is the argument, and it is the reason this post exists.

Every video model in this series has the same fundamental weakness, and every vendor admits it in their own words. ByteDance name the physical plausibility of complex motions. Alibaba name consistency across shots. Google’s parallax is convincing rather than correct. The models are extraordinary at what a thing looks like and approximate about what a thing is: how it weighs, how it resists, how it falls, what happens when a hand actually closes around it.

This is not a resolution problem and it will not be fixed by more pixels. It is a physics problem. The models learned motion by watching video, which means they learned the appearance of physics without ever learning physics.

A model trained jointly on robot action prediction is being forced, in a way no purely generative model ever has been, to represent consequence. An arm that predicts wrongly about a soft material does not produce an implausible frame. It drops the part. The feedback is physical and unforgiving, and it has to be right within about a tenth of a second.

Whether that representation actually transfers to the video tier is an open question and we are not going to claim it does. BFL’s architecture says it should, since joint training across modalities is the entire premise of Self-Flow. What we can say is that this is the first serious attempt anyone has made to teach a generative video model about the physical world through something other than more video, and the weakness it targets is precisely the weakness every other vendor is currently apologising for.

It is also, if you want it, a rather beautiful idea: that the way to make a shot of a hand picking up a cup look right is to first teach the machine to actually pick up the cup.

What it changes on the floor

Native multilingual dialogue is the feature Indian and other multilingual production has been waiting for. Not audio, but dialogue. Generated in the same pass as the image, synchronised, in a language you choose. Anyone who has worked in a market where one film ships in eight languages understands what a post-production line item that represents. We will hold judgement until we have run it across Indic languages at volume, because “multilingual” has historically meant the six big ones, but the direction is right.

Keyframe-to-video is the storyboard, again. Like Omni’s first-and-last-frame control, this is a model accepting the instruction format directors have always used. The convergence across vendors on this idea in a single quarter is not a coincidence. It is the industry discovering that prose is a poor way to specify a camera.

Agentic chaining for multi-shot is a trade-off rather than a feature. Chaining means each shot is generated with awareness of the last rather than as part of a single continuous pass. That gives you more control per shot and more opportunity for drift across shots. Know which one you are buying. For a sequence with deliberate cuts, chaining is arguably the more honest mechanism.

Twenty seconds sits mid-field. Longer than H3’s fifteen, shorter than Seedance and Wan’s thirty.

The grammar test

For FLUX 3 Video we run the weight test, chosen deliberately to target the physics claim.

Generate three shots, each five to eight seconds, each about a single physical event.

  1. A hand setting a full glass down on a wooden table. Watch the moment of contact. Does the glass decelerate before it lands, or does it arrive? Does the liquid acknowledge the stop?
  2. A length of cloth dropped from frame-top onto a chair. Watch how it settles. Cloth is the classic tell, because models produce fabric that behaves like a cloud, drifting to rest without ever having had mass.
  3. A person changing direction at a walk. Watch the weight shift. Bodies lean before they turn. Generated bodies usually turn and then, unconvincingly, lean.

Score each on a single question: did the object have mass?

Everything else, including resolution, grade and prompt adherence, is a spec conversation. This is a craft conversation, and it is the one that decides whether an audience believes what they are watching. Audiences cannot articulate why a shot of a weightless glass feels wrong. They feel it instantly and completely.

FLUX 3 Video, in our runs, performs above the field on test one, comparably on test three, and is still clearly solving an appearance problem on test two. Cloth remains unsolved by everybody. We will re-run this test when FLUX 3 Dev ships and the action training has had another generation to mature. The weight test is the one benchmark in this series that we expect to actually move as a result of architecture rather than scale.

Where it breaks

  • Half the model is a waiting room. Image has no public endpoint well past “coming weeks.” Dev has no date and no licence terms. Action is partner-only. Plan for Video and treat the rest as roadmap.
  • Benchmarks are BFL’s own, on an early candidate, at 720p, on 10-second clips, with no independent replication. Read them that way.
  • Full resolution capability undisclosed. Benchmarks ran at 720p. What the model does above that was not published at announcement.
  • Agentic chaining drifts. It is a different failure mode from single-pass extension rather than a better one.
  • Twenty seconds will feel short if you have been working at thirty.
  • Cloth still floats. Everyone’s does.

Inside Pindow

Canvas, with a caveat we apply to every gated model.

Our rule on early access is simple and we would recommend it to anyone building a production pipeline. No model enters the working chain until it has a public endpoint and a stable model ID. Not because gated models are not good, since FLUX 3 Video is genuinely strong, but because a pipeline built on an access list is a pipeline that can be revoked. FLUX 3 Video cleared that bar in August. FLUX 3 Image has not cleared it yet, and until it does it is not in the graph.

That is not a criticism of BFL. Shipping a four-modality unified architecture is hard and they were honest about sequencing from day one. It is a description of how a platform that people depend on has to behave.

When FLUX 3 Dev arrives with open weights, it joins the same archival logic we described for H3: pin the checkpoint, keep the look. We are watching the licence terms closely, because an open-weight multimodal backbone with a workable licence would be the most significant thing to happen to independent AI filmmaking in two years, and an open-weight backbone with a restrictive one would be a press release.

Working notes

Test for weight before you test for beauty. A gorgeous shot in which nothing has mass will not cut against a plain shot in which everything does.

Use keyframe-to-video for anything with a move. Two frames beat any sentence.

Treat chained shots as a sequence rather than a take. Check continuity at every join, the way you would across a cut, because that is what it is.

Do not build on gated access. Prototype on it, by all means. Do not promise a client a delivery date that depends on somebody approving your application.

Run your own comparison before believing any preference number, including ours. Vendor benchmarks measure what the vendor chose to measure. Your shot list is the only benchmark that matters.

The most interesting sentence in the FLUX 3 announcement is not about video at all. It is the one about a robot arm learning to handle soft material in a hundred milliseconds, because the thing that has always separated a convincing shot from an impressive one is not how it looks. It is whether the world in it has weight.

Just make a film.

Sources

FLUX 3: Multimodal Video, Image & Audio (Black Forest Labs) · Black Forest Labs launches FLUX 3 (VentureBeat) · Black Forest Labs Releases FLUX 3 (MarkTechPost)

Related Stories

View all →