The Layout Problem Is Solved. The Taste Problem Isn’t.
Model GPT Image 2.5 (Flare · Sunburst) · Launched 8 September 2026 · Team Pindow · 9 min

GPT Image 2.5 can finally hold a dense frame together. That makes the art director’s job harder, not easier.
For three years, the honest limitation of every text-to-image model was the same, and nobody wanted to say it plainly. They could make a beautiful picture. They could not make a designed one. Put four elements in a frame with a hierarchy between them, a headline and a product and a body of text and a logo lockup, and the model would give you something that looked like design from six feet away and fell apart the moment you leaned in. Letterforms dissolved. Kerning invented itself. The third element in the hierarchy quietly became the first.
So we worked around it. We generated the image and set the type in a real layout tool, the way you would shoot a plate and do the titles in post. That workaround was correct. It is now, for a large class of work, optional.
On 8 September, OpenAI shipped GPT Image 2.5. The release note leads with speed. That is not the story. The story is that dense layouts stay readable.
What actually shipped
Two models rather than one, and the split matters more than the naming suggests.
GPT-Image-2.5 Flare is the default. OpenAI’s framing is higher quality than GPT Image 2 at roughly half the latency, with generation overall around 50 percent faster than the 2.0 generation. This is the workhorse: boards, variants, the two hundred frames you throw away before you find the one.
GPT-Image-2.5 Sunburst is the precision tier, built for detail that has to survive scale. Images that will be inspected closely or printed large. Same API shape as Flare. You choose by intent.
The shared capabilities:
- Resolution up to 3840px on the long edge. Presets for square, portrait and landscape, with custom dimensions accepted in multiples of 16.
- Six quality tiers, from auto through to max. The top two are new headroom above anything 2.0 could reach.
- Up to 16 reference images per edit request. This is the number that changes workflows.
- Transparent backgrounds, delivered as PNG or WebP. Native alpha rather than a matte you have to pull.
Alongside the models, four product-side additions. A sketch tool that lets you draw a rough directly in ChatGPT as visual reference. Templates for recurring formats like posters and product shots. Prompt sharing. And in-image comments, where you annotate a region and ask for a change there specifically.
It went live across ChatGPT, ChatGPT Work and Codex on every tier, on desktop, mobile and web, with both API model IDs callable the same day. That last detail is worth noting for anyone who has spent 2026 on early-access waitlists. Announced and shippable arrived together.
What it changes on the floor
The key art round gets shorter by a day. Not because the model is faster, but because the layout comes back intact. When a title card generates with the type legible and the hierarchy holding, the client conversation moves from “the text is broken, ignore it” to “the text says the wrong thing.” That is a different meeting. It is the meeting you actually wanted.
Sixteen references is a character bible rather than a mood board. Twelve months ago, a reference-driven edit meant one image and a prayer. Sixteen slots is enough to carry a face from three angles, a costume from two, a location plate, a lighting reference, a texture swatch and a grade target, and still leave room for the thing you forgot. For anyone building a series rather than a single frame, this is the difference between re-establishing your character every session and simply continuing.
Native alpha changes compositing order. A transparent PNG straight out of the generator means the element arrives ready to sit in a layout or a comp rather than needing a key. Small thing. Saves an hour per asset, every asset.
In-image comments are, quietly, the most cinematic feature in the release. Pointing at a region of the frame and saying this, not that is how notes have been given on set and in the edit suite forever. It is closer to a grease pencil on a monitor than to a prompt box. Any feature that moves AI direction back toward the physical grammar of giving a note is a feature we pay attention to.
The new top quality tiers reintroduce a real gradient. For most of the last two years, the difference between a draft and a finish was negligible enough to ignore. It is not now. A 4K maximum-quality plate is a meaningfully heavier operation than a 1024 draft, which means the old discipline returns: draft rough, finish properly. Anyone who came up shooting knows this instinct already. You do not roll the expensive stock on the rehearsal.
The grammar test
Every Model Note runs one fixed directorial exercise. For image models we run the three-element hierarchy.
Ask for a single frame containing exactly three readable elements at three levels of dominance: a face in the near ground, a legible printed sign in the mid ground, and a distant practical light source that motivates the key. Specify which of the three the eye should reach first, second and third. Do not describe style. Do not describe mood. Describe only the hierarchy.
Then look at what came back and ask three questions.
- Did the eye actually travel in the order you specified?
- Is the sign readable, not “texty” but readable, and in the language you asked for?
- Does the distant practical actually motivate the light on the face, or is the face lit by nothing?
Most models pass the first, fail the second, and have never once passed the third. GPT Image 2.5 passes one and two reliably in our runs. Three remains inconsistent. It will give you a beautiful frame in which the key light has no source in the world of the picture, and it will do it confidently.
That third question is not a nitpick. Motivated light is the difference between an image and a shot. A frame where the light has no origin can be lovely and will still cut badly against a frame where it does, because the audience reads continuity of light before it reads anything else. This is the oldest grammar there is, and it remains the thing generative models understand least.
Where it breaks
- Light without a source. As above. The model composes light. It does not yet reason about where light comes from. Specify the source, its colour temperature and its direction explicitly, every time, or accept whatever it invents.
- Non-Latin typography is still uneven. Devanagari, Urdu and most Indic scripts render better than a year ago and still require a designer’s eye before anything goes to a client. Treat generated Devanagari as a comp rather than final artwork. Set the real type.
- Multi-turn drift is reduced rather than eliminated. Consistency across an editing chain is visibly better than 2.0. It is not free. Past roughly six or seven turns on the same asset you will still see the face soften toward an average. Re-anchor with your references rather than trusting the chain.
- The top tiers are yours to manage. Nobody will stop you running maximum quality at 4K on a thumbnail.
Inside Pindow
This one lands squarely in Canvas.
Our position on image models has not changed and will not. The model is the lens, not the eye. Pindow’s job is to hold the decisions that belong to the director, meaning style, reference, consistency and continuity, steady above whichever model is currently best at making pixels. When a new image model ships, the thousand-plus styles in the Canvas library and the reference structure around them do not get rebuilt. They get pointed at something new.
Practically, on the Canvas node graph: Flare for exploration, Sunburst for the frame you are going to live with. Generate wide while you are still deciding. Switch tiers at the point where you stop asking what if and start asking is this it. The Prompt Engine handles the syntax difference between the two, so you change the node rather than rewriting the prompt.
The 16-reference ceiling maps cleanly onto Character Studio. A locked character sheet plus a location plate plus a grade target fits comfortably inside the budget with slots to spare, which means consistency stops being something you fight for each session and becomes something the graph carries for you.
Working notes
Name the light source before you name the mood. “Golden hour” is a mood. “Low sun, camera left, 40 degrees off axis, warm bounce from a whitewashed wall camera right” is a shot. The model rewards the second and guesses at the first.
Use the reference budget structurally rather than greedily. Sixteen slots filled with sixteen versions of the same face is worse than six slots covering angle, costume, location, lighting, texture and grade. References should span the axes of variation you care about rather than piling up on one.
Draft small, finish at your delivery size. Not because the model cannot do 4K early, but because deciding should feel cheap and committing should feel expensive. That instinct is worth protecting.
Treat generated type as a comp. Always. The layout can be trusted now. The typography still belongs to a designer.
The layout problem was a technical problem, and technical problems get solved on somebody else’s schedule. What the model cannot do is decide which of the three elements deserves the eye first. That decision was never a spec. It was always taste, and taste is still built the slow way, by looking at a great deal of work and understanding why it holds.
Just make a film.






