Invest2 publishers2 min readPublished
OpenAI's Images 2.5 ties Nano Banana 2 at three categories each in Decrypt's rerun
Decrypt's rerun gives three of six categories to Google and the rest to OpenAI. The 50% latency cut is OpenAI's own figure, measured against Images 2.0, and the review does not include prices for the two new API tiers.
The Investor · Invest desk

What happened
- OpenAI shipped ChatGPT Images 2.5 on September 8, claiming sharper detail, more precise editing and up to 50% lower latency than its predecessor.
- Two models went into the API alongside it: GPT-Image-2.5 Flare as the fast default and GPT-Image-2.5 Sunburst built for premium editing precision.
- In the text-heavy street scene, Images 2.5 rendered its own graffiti tag with an extra L and lost the apostrophe in KELLERMAN'S, and Decrypt handed that category to Nano Banana 2.
- The Google entrant was Nano Banana 2, its name for Gemini 3.1 Flash Image, with the slower Nano Banana Pro left out so the two fast models met on the same tier.
Compiled by The InvestorSomething wrong?How this is made
Why it matters
- decision With the categories split evenly, the pick turns on which defect a given job can absorb, so a poster shop and a mood-board tool inside the same company will reasonably choose different models.
- constraint OpenAI's 50% latency figure is measured against its own Images 2.0, so any team comparing Flare with Gemini 3.1 Flash Image has to time both itself before it can budget throughput.
- exposure A team standardising on OpenAI for customer-facing text carries the per-generation defect risk: the yellow cast and the oversharpening each cleared only when the next model shipped.
- capability Sketch, prompt sharing, inline comments on image regions and poster and merch templates change what a design team can do around the model, and a side-by-side image test scores none of it.
In May the same evaluator ran eight categories, and GPT Image 2 took more of them than Nano Banana 2 did [7]. No ties were reported, so "more of eight" means five at minimum, or 62.5%; three of six is 50% [9]. OpenAI's share of scored categories went down, in other words, in the release whose pitch is precision. The rounds are not strictly comparable, since the six categories were carried over or adapted from the earlier ones and fed as identical prompts to both models [8].
The 50% is OpenAI's own figure, and it compares Images 2.5 with Images 2.0 [2]. What Decrypt scored is image content, prompt by prompt [4]. The review does not include prices for Flare, Sunburst, or the new high and max quality tiers that now sit above where Images 2.0 topped out [17][16]. Until those land, a three-three split cannot be converted into a cost per acceptable image, and a cost per acceptable image is what an operator running thousands of poster variants actually needs.
Split by workload, the two models fail differently. For output a customer reads, what breaks is legibility, and Decrypt's text-heavy street scene put two slips on Images 2.5 against a single garbled payphone sticker on Google's side [13][12]. For output where atmosphere carries the job, Images 2.5 took the steampunk aerial by a clear margin [14].
OpenAI's image releases have run one signature defect per generation. GPT Image 1 carried a warm yellow cast the company never fully explained and never fully fixed [10]. GPT Image 2 traded it for oversharpening whenever a prompt stacked too many constraints [7]. This round, color balance and detail held at full complexity, and Decrypt reports neither defect [11].
Two legibility slips in one image is thin evidence for a third defect in that sequence, and the way to settle it is volume: if misspelled signage recurs across a few hundred prompts, the fix was a swap. I'd expect Google to hold the edge on anything with words in it, and the case against me is in Decrypt's own summary, which puts the swing between the two down to a few checkable mistakes instead of very noticeable issues [5].
What to watch
- Published per-image prices for Flare and Sunburst at the new high and max tiers. Those prices would let the three-three split be costed.
- Whether misspelled or garbled text recurs for Images 2.5 across a larger prompt set. If it recurs, that is the third generation's signature defect.
- A rerun against Nano Banana Pro, the slower Google model Decrypt left out of this same-tier test.