The Model You Can Take Home
Model MiniMax H3 (Hailuo 3.0) · Launched 31 July 2026 · Team Pindow · 10 min

H3 is fifteen seconds at 2K with native stereo audio. It is also, days after launch, sitting on Hugging Face with its weights open, and that second fact is the one that will still matter in three years.
Two models shipped on 31 July 2026. One of them, Seedance 2.5, got the coverage, and reasonably so, since it is the larger change to a narrative shot list.
The other one gave itself away.
MiniMax released H3, also known as Hailuo 3.0, and then published the weights. Not a distilled version. Not a research preview with a non-commercial licence and a shrug. The actual omni-modal model, downloadable, running locally, with community tooling support arriving within days.
We want to be careful not to romanticise this, because open weights are not automatically better weights and a model you can run is not automatically a model you should. But there is a specific reason this matters to the kind of filmmaker we build for, and it is not ideology. It is permanence.
Every hosted model in this series can be deprecated. Gemini’s preview endpoint dies on 30 September. Model IDs get retired. Access terms change. Regions get cut off, and Seedance’s official international route already excludes several English-speaking markets outright. If your film’s look depends on a hosted checkpoint, your film’s look has an expiry date set by somebody else’s product roadmap.
A film shot on a particular stock keeps looking like that stock forever. That property has been missing from generative work since it began. H3 is the first model in this quarter’s releases where you can have it back.
What actually shipped
MiniMax H3 (Hailuo 3.0), released 31 July 2026, with weights published on Hugging Face shortly after and community node support, including reference-to-video and first-and-last-frame workflows, landing within the first week of August.
Output. What comes back.
- Video up to 2K resolution, maximum 15 seconds.
- Native stereo audio with every generation. Not a separate pass, not a mono bed, but stereo generated jointly.
Inputs, which is the “omni” part. What it will accept.
- Text, images, video and audio all accepted as context, in the same context window.
- Image, audio and video reference and editing.
- Video-to-video motion transfer.
- Multi-shot native modelling, meaning the model handles shot changes as a modelled phenomenon rather than as an accident.
- Instruction following and text rendering explicitly called out as improvements.
Open weights. MiniMax stated at launch an intent to open the weights “in the coming days, subject to applicable laws and regulations,” and followed through. The model card is published under MiniMaxAI/MiniMax-H3.
Vendor-stated limits. MiniMax say H3 “still has room to grow,” with planned improvements to multimodal understanding and visual detail. Fifteen seconds is short against a field where thirty is becoming standard, and 2K sits below the native-4K claims elsewhere.
What it changes on the floor
Stereo, generated jointly, is a bigger deal than the spec line suggests. Mono generated audio has always been a temp track. You could use it for timing and nothing else, because it carried no spatial information. Stereo generated in the same pass as the image means the audio knows where things are in the frame. A door closing camera left sounds like it closed camera left. That does not make it a deliverable stem, because it is not, but it makes the offline cut readable in a way mono never was, and anyone who has tried to judge the rhythm of a sequence against mono scratch knows how much that costs you.
Fifteen seconds is a lighter unit, and that changes what you are willing to throw away. The economics of exploration are the economics of craft. If a take is a lighter commitment, you take more of them, and taking more of them is how you find the one nobody planned. Light does not mean disposable. It means you can afford to be wrong on the way to being right.
Motion transfer is direction by demonstration. Video-to-video motion transfer means you can show the model the move instead of describing it. Shoot ten seconds of a handheld follow on your phone, transfer the motion, keep your subject. For anyone whose instinct for camera lives in their hands rather than in their vocabulary, this is a more natural instruction set than any prompt. It is also, practically, how you get the specific unsteadiness of a real shoulder-mounted shot, which no adjective has ever successfully described.
Multi-shot as a modelled phenomenon rather than an accident. Most video models produce a shot change when they lose the thread. A model that treats a cut as something it is doing rather than something that happened to it is a model you can direct across a cut.
Open weights mean a locked look. Download the checkpoint that made your film. Archive it with the project. In five years, when you need one more shot for a re-cut, the model still exists, on your drive, at your version, behaving identically. Nobody in this industry has been able to say that before, and it is an unglamorous, load-bearing kind of freedom.
The grammar test
For H3 we run the archive test, because it is the only model in this quarter where the test is even possible.
- Generate a shot you are happy with. Record everything: prompt, references, seed, parameters, and the exact model checkpoint.
- Archive the checkpoint alongside the project files.
- Wait. A week is enough to prove the mechanism. The real test is a year.
- Regenerate from the archive, locally. Compare frame by frame.
A hosted model will fail this test eventually and without announcement, because the checkpoint behind the endpoint is not a contract. A local checkpoint passes it by definition.
Then run the second half, which is the part that actually tests craft.
Generate a six-shot sequence, running establishing, wide, medium, close, reverse and out, as six separate 15-second generations from a single locked character and location reference set. Cut them together and watch for the three things that betray AI sequences.
- Does the light stay on the same side across the reverse? Reverses are where AI sequences die. The model has no idea there is a line.
- Does the character’s face survive six generations? Drift shows up worst between the wide and the close, where the model has the most freedom to reinterpret.
- Does the stereo field stay oriented? If the traffic was left in shot one, is it left in shot four?
H3 holds one and two well for a 15-second model with a disciplined reference set. Three is genuinely impressive and genuinely inconsistent. When it holds, the sequence has a spatial coherence we have not heard from generated audio before. When it does not, it flips without warning, which is worse than mono.
Where it breaks
- Fifteen seconds is short. In a quarter where 30 is becoming the number, H3 asks you to think in shorter units. For a lot of grammar, like a held look or a slow reveal, that is a real constraint.
- 2K is not 4K. For broadcast and cinema delivery you are upscaling. For digital and social you are fine.
- Stereo orientation flips. Sometimes. Check it every time before you trust the offline.
- Running locally is real work. Open weights are not a consumer feature. Compute, VRAM, and a node graph somebody has to maintain. The freedom is real and it is not free.
- Visual detail has room to grow, per MiniMax. Fine texture at 2K is where you will see it.
- Open weights are a responsibility. A model anybody can run is a model anybody can run. We are not going to pretend that question does not exist, and we would rather the industry discussed it soberly than performed either panic or indifference.
Inside Pindow
Canvas, and H3 occupies a particular position in how we think about the platform’s job.
Pindow aggregates best-in-class tools across the workflow. The unglamorous truth underneath that sentence is that best-in-class changes constantly, and a platform whose value is “we picked the best model” is a platform with a shelf life of about four months. Our value has never been the picking. It is the layer above: the two hundred-plus filmmaker presets, the Prompt Engine, the style library, the reference structure, the consistency machinery. That layer is model-agnostic on purpose, because the model underneath it is the one thing guaranteed to be replaced.
H3 is the clearest illustration of why we build that way. It is a model that can be pinned. A director who finds a look on H3 and archives the checkpoint has made a decision that survives every subsequent product announcement, and Pindow’s job is to make that decision portable, keeping the grammar, the presets and the structure identical whether you are calling a hosted endpoint or a local checkpoint on your own machine.
The motion-transfer path wires naturally to the 3D Stage as an alternative route. Block spatially when you want geometric precision. Transfer motion when you want the feel of a real operator’s hands. Those are two different kinds of truth and a director should be able to choose.
Working notes
Archive the checkpoint with the project. Every time. It costs disk and buys permanence. This is the single most valuable habit available to anyone working with open-weight models and almost nobody does it.
Design for fifteen seconds. Do not fight the ceiling. Use it. A 15-second limit is roughly a shot, and thinking in shots rather than in durations is the better habit anyway.
Check stereo orientation on every take. It is a five-second listen and it will save you an assembly.
Shoot your own motion references. A phone, ten seconds, your own hands. The unsteadiness you get from a real operator is not available in any preset.
Use the lighter unit on exploration rather than on volume. More attempts at the right shot, not more shots. The temptation runs the other way.
The models that will be remembered from 2026 are probably not the ones with the best numbers. They are the ones somebody could still run in 2031, because a filmmaker archived the weights alongside the project the way an earlier generation kept a can of the same emulsion batch in the fridge.
Just make a film.






