←Back to News
FILMMAKING—Aug 24, 2026—10 Min Read

The Deck Is Not a Film

Model Wan 3.0 · Launched 24 August 2026 · Team Pindow · 10 min

Pindow and Wan 3.0

Wan 3.0 will turn a fifty-page PDF into thirty seconds of video. That capability is genuinely useful and genuinely dangerous, and the difference is entirely a matter of craft.

Alibaba shipped Wan 3.0 on 24 August, a few days after a large share sale, and the headline feature everybody ran with was document-to-video. Feed it a deck, get a film.

We want to sit with that for a moment, because the reflex in our industry is to treat a feature like this as either a miracle or an insult, and it is neither.

Here is the useful version. A training department has 200 slides of safety procedure that nobody reads. Turning that into watchable video used to require a budget nobody was going to approve, so it stayed as slides, and people kept not reading it. That is a real problem and this is a real solution to it.

Here is the dangerous version. A brand has a strategy deck. Somebody feeds it to a model. Thirty seconds come out. Everyone agrees it looks quite good. Nobody notices that the film has no point of view, because a deck has no point of view. A deck is a list of true statements, and a film is an argument with an emotional shape. The model has faithfully converted the first into the second’s clothing.

The feature is not the problem. The absence of a director is the problem. Let us look at what actually shipped.

What actually shipped

Wan 3.0 launched officially on 24 August 2026 after a public beta earlier that month, available through Alibaba Cloud Model Studio.

  • Duration of up to 30 seconds per generation, with smart duration recommendation and a video-extension feature for longer pieces.
  • Resolution at 480p, 720p and 1080p.
  • Document input accepting doc, xls, ppt, pdf, txt, key, pages, numbers and md. One file or link per request, up to 100MB and 50 pages.
  • Multimodal input beyond documents, covering text, images, audio and video.
  • Reference-to-video mode maintaining consistency across characters, props and spatial relationships.
  • Video editing without regenerating from scratch.
  • Faces. Alibaba make a specific point of “diverse, lifelike human faces” rather than the repetitive AI-generic look, and having run it, we think the claim holds up better than most claims of this kind.
  • Improvements cited at launch in instruction following, cross-shot consistency and audio quality.

Alibaba are unusually candid about the limits in their own materials. Audio texture and on-screen text rendering accuracy are, in their words, still improving and not yet where they want them. We would rather a vendor said that than did not.

What it changes on the floor

A moving-image pitch stops being a budget question. This is the change that matters and nobody is saying it plainly. A thirty-second 1080p generation is now something a director can produce in an afternoon, which means arriving at a client meeting with the film instead of the board. Not a good film. A rough, honest, thirty-second articulation of what the idea feels like. For anyone who has spent fifteen years watching great ideas die in a room because the client could not picture them, this is not a workflow improvement. It is a structural change in who gets to sell an idea.

We should be clear about who this benefits. It is not the agency with a previz department. It is the person who never had one.

The face diversity claim deserves attention. For three years, generative faces have converged on an oddly specific global average: a symmetrical, mid-twenties, softly-lit non-specificity that shows up whether you asked for a Punjabi farmer or a Warsaw tram driver. It is the visual equivalent of a stock photo. If Wan 3.0 genuinely holds specificity of face better, that is worth more to the kind of work we care about than any resolution bump. A film about real people needs faces that look like they came from somewhere.

Document input is a research tool before it is a generation tool. The interesting use is not “deck in, film out.” It is “here is a fifty-page brand guideline document, here is a scene, make sure the scene does not violate the document.” Using the ingest path to give the model context rather than to give it the script is a much better idea than the one in the press release.

Editing without full regeneration is the quiet productivity feature. Anyone who has re-rolled an entire thirty-second generation to fix one second knows exactly what this saves.

The grammar test

For document-to-video we run the deck test, and it is deliberately unkind.

Take a real brand deck of ten slides, with strategy language, bullet points and a tagline at the end. Feed it in. Take whatever comes back and ask three questions of it.

  1. Where is the turn? Every piece of film that works has a moment where the thing changes: a reveal, a reversal, a shift in understanding. Find it in the output. Point at the timecode.
  2. Who is the film about? Not what it is about. Who. Name the person on screen whose want drives the thirty seconds.
  3. What does the audience feel at second 28 that they did not feel at second 2?

In our runs, the output reliably fails all three, and we want to be precise about why, because the failure is not the model’s fault.

It fails because the deck did not contain the answers. A strategy document lists propositions. It does not contain a protagonist, a want, an obstacle or a turn, because those are not what a deck is for. The model performed a faithful conversion of the input. The input simply was not a film.

Run the same test with a properly structured one-page treatment, with a protagonist, want, obstacle, turn and resolution, and the output passes two of the three questions comfortably. Same model. Same thirty seconds. Entirely different result, because somebody did the writing.

That is the whole lesson of document-to-video, and it is an old lesson wearing new clothes. The tool will execute your structure faithfully, including when you do not have one.

Where it breaks

  • On-screen text rendering is, per Alibaba’s own note, not yet reliable. Do not put your end-card typography through the model. Set it properly.
  • Audio texture is similarly flagged by the vendor. Treat generated audio as scratch.
  • The 50-page and 100MB ingest ceiling is generous but real, and long documents get summarised aggressively. What survives the summary is not under your control.
  • 30 seconds is still a shot sequence rather than a story. Extension exists, but chaining extends accumulates drift like every other model in this series.
  • Structure in, structure out. The most important failure mode is not technical. Feed it a list, get a list.

Inside Pindow

This one touches two stages, and the order matters. Script first, then Canvas. Never Canvas first.

The Pindow position on document ingest is that it belongs at the Script stage rather than the generation stage. A brand document, a brief, a research pack, a set of guidelines: these are inputs to writing, and writing is where structure gets decided. Fountain import and the Catalogue exist so that a brief becomes a screenplay with characters and locations as real objects, and then becomes shots. Skipping the middle step does not save time. It relocates the failure to somewhere more expensive.

So in Pindow, Wan’s document capability sits behind the Script stage as a context loader. Load the brand guidelines. Load the research. Let them inform the draft. Then write the treatment, with protagonist, want, obstacle and turn, and only then push shots into Canvas.

On the Canvas side, the practical value is the resolution ladder. 480p for blocking, 720p for the internal review, 1080p only for what you would actually show. Blocking is light enough that there is no excuse for showing a client the first thing that came out.

One more thing, and it is the one we feel most strongly about. The thirty-second pitch film made in an afternoon is the most genuinely democratising thing in this entire quarter’s releases. It puts a moving-image pitch within reach of somebody with a laptop and an idea, in a market where that has never been true. The antagonist in this story has never been the technology. It has always been access: who gets the previz budget, who gets the room, whose idea gets to be seen before it gets judged. Anything that widens that door is on the right side.

Working notes

Write the one-pager before you touch the ingest. Protagonist, want, obstacle, turn, resolution. Five lines. If you cannot write them, the model cannot find them.

Use document ingest for context rather than for script. Guidelines in. Screenplay out of your own head.

Block at 480p. Always. It is cheaper than thinking about it.

Set your own type. Every time. The vendor told you it is not ready. Believe them.

Cast the face deliberately. Wan holds facial specificity better than most. That is an invitation to describe a real person, with age, region, weathering and the particular tiredness of a particular job, rather than to accept the average. Specificity of place and character is not a style preference. It is the difference between a film about people and a film about nobody.

A deck is a list of things that are true. A film is one thing that is felt. Any tool that converts the first into the shape of the second is doing you an enormous favour and a small, quiet violence at the same time, and the only defence is to have written the film before you asked the machine for it.

Just make a film.

Sources

Wan3.0: 30-Second AI Video Generation from Any Input (Alibaba Cloud) · Alibaba launches Wan3.0 video model with 30-second generation and document input (TechNode) · Alibaba launches Wan3.0, its 30-second video model (TNW) · Alibaba Launches Wan3.0 AI Video Model (WinBuzzer)

Related Stories

View all →