LeadershipNot yet confirmed elsewhere1 publisher3 min readPublished
Lower-priced AI models ran up bigger bills in 32% of head-to-head tests
Researchers led by Lingjiao Chen of Stanford found the AI model with the lower list price cost more in 106 of 336 head-to-head comparisons. For anyone setting an AI budget, the finding means per-task cost has to be measured on their own workloads before a model is chosen.
The Board Room · Leadership desk

What happened
- At May 1, 2026 list prices, Gemini 3 Flash cost 80% less than GPT-5.4, but across the study's tasks it came out 38% more expensive.
- On one MMLU-Pro problem, Gemini 3 Flash wrote more than 60,000 hidden thinking tokens, while GPT-5.4 solved the same problem with 25.
- Running the same query again and again on one model, the costliest run cost up to 9.7 times as much as the cheapest.
- Uber had exhausted its entire 2026 AI budget by about four months into the year.
- Only 11% of 396 enterprises surveyed by Benchmarkit and Mavvrik in April and May could forecast AI costs within 10%, down from 15% in 2025.
Compiled by The Board RoomSomething wrong?How this is made
Why it matters
- decision Model selection now needs a costed trial on a team's own tasks before signing, and the researchers' published per-run data and code make that trial cheaper to set up.
- constraint Run-to-run variance survives re-prompting, so measured usage gives a cost range per query and fixed per-task chargebacks to internal teams get harder to set.
- cost In agent workloads a cheap model that loops is billed for every turn, so the bill for a failed run can exceed a pricier model's bill for a finished one.
A rate card prices tokens. What a task costs depends on how many tokens the model decides to spend on it, and that count varies between models and between runs of the same model [20]. For Gemini 3 Flash's 80% discount to finish as a 38% premium over GPT-5.4, Flash had to bill about 6.9 times the token volume, assuming the discount held evenly across input and output [23]. The researchers traced much of that volume to thinking tokens, the hidden reasoning a model writes before it answers [18].
Agent work widens the gap. When a model works through tools or a computer environment, each turn can carry the earlier conversation back in as input, and in those tasks the number of turns drove cost [12]. On one prompt, Gemini 3.1 Pro finished in 85 steps for about $1. Gemini 3 Flash took nearly 1,000 steps, ran up $14 in token charges and failed [13]. That is roughly 12 times the steps for 14 times the money [24].
A finance lead could fairly read the headline result the other way, since the cheaper-listed model did not cost more in 230 of the 336 comparisons [22]. The 336 matches every pairing of the eight models on each of the 12 tasks [8][21], so each reversal belongs to one pair on one task. No model was consistently the cheapest or the most expensive; the ranking changed from task to task [9]. A buyer who knows the rate card was right about two times in three still cannot tell which of their own tasks sit in the other third.
Measuring on a team's own tasks helps more with choosing a model than with fixing a budget. The paper describes an irreducible noise floor: models can follow different reasoning paths even when their inputs stay fixed, and re-prompting does not remove the variation [2]. In a follow-up on two programming prompts, models from Anthropic, Google and OpenAI each got the same prompt five times and produced a different cost each time [14]. "The practical takeaway is clear," said Lingjiao Chen, a researcher at Stanford University and Microsoft Research [7]. "Price alone should not be used to infer which model is actually cheaper." [15] The team published its per-run cost data and code so companies can repeat the comparison on their own workloads [16].
Evidence that companies are missing their forecasts is of uneven quality. The survey comes from a vendor: Mavvrik sells AI cost-management software, and the figures are self-reported [19]. Uber's is a single company's account. Its chief technology officer, Praveen Neppalli Naga, used $1,200 of tokens in a two-hour demonstration of the company's coding tools in April [3], about $600 an hour [25]. "I'm back to the drawing board, because the budget I thought I would need is blown away already," Naga said [5]. The account of Uber's exhausted budget does not say how much of the overrun came from model choice and how much from volume [4].
The trade-off this quarter is between paying for a measurement run and picking from the rate card. The run costs tokens and engineering time now. Each task has to be repeated enough times to show the spread between cheap and expensive runs. I think the output worth paying for is a cost range per task type with its tail included, since that range is what next quarter's budget has to absorb. A team that skips the run sets that budget from a price comparison that reversed in 106 of 336 tests [6].
What to watch
- Replications that run the researchers' published per-run data and code against production workloads, and whether their reversal rates land near the study's 32%.
- Any explanation from Uber of what drove its 2026 AI budget overrun and how it reset the budget.
- Provider moves to cap or itemise thinking tokens on bills, the volume the study found helped explain the reversals.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap+8
- Incentives45
- Confidence58
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
In the May revision, repeated runs of an identical query on the same model varied by up to 9.7 times between the cheapest and most expensive run.
ReportedSupportedSource: implicator.ai2 sources— create a free account to open themView cited source - [2]
The paper describes an irreducible noise floor: models can follow different reasoning paths even when inputs stay fixed, and the variation cannot be removed by re-prompting.
ReportedSupportedSource: implicator.ai2 sources— create a free account to open themView cited source - [3]
In April 2026, Uber chief technology officer Praveen Neppalli Naga demonstrated the company's AI coding tools and over two hours used $1,200 worth of tokens.
- [4]
By April 2026 Uber had already exhausted its entire 2026 AI budget, about four months into the year.
- [5]
"I'm back to the drawing board, because the budget I thought I would need is blown away already," Naga said.
- [6]
Researchers from Stanford, Carnegie Mellon, UC Berkeley and Microsoft Research found that in 106 of 336 pairwise comparisons (32%), the model with the lower listed price cost more in total.
ReportedSupportedSource: Price Reversal Phenomenon study, reported by implicator.aiView cited source - [7]
Lingjiao Chen, a researcher at Stanford University and Microsoft Research, tested whether listed API prices predict actual running cost, with co-authors from Carnegie Mellon, UC Berkeley and Microsoft Research; the paper, The Price Reversal Phenomenon, first appeared March 25, 2026 and was revised May 28.
- [8]
The revised study tested eight frontier reasoning models across 12 tasks.
- [9]
No model in the study was consistently the cheapest or the most expensive across its benchmarks; the ranking changed from task to task. Listed price was defined as input plus output rates.
- [10]
Gemini 3 Flash was listed 80% cheaper than GPT-5.4 at May 1, 2026 prices but cost 38% more across the study's tasks.
- [11]
On one MMLU-Pro problem in the study, Gemini 3 Flash consumed more than 60,000 thinking tokens; GPT-5.4 solved the same problem with 25.
- [12]
For tasks requiring a model to interact repeatedly with tools or a computer environment, the number of turns also drove costs; each turn can include earlier conversation history as input.
- [13]
On one prompt, Gemini 3.1 Pro finished in 85 steps for about $1, while Gemini 3 Flash went through nearly 1,000 steps, accumulated $14 in token charges and failed.
- [14]
A follow-up analysis on two selected programming prompts gave models from Anthropic, Google and OpenAI the same prompt five times each, with a different cost on each run.
- [15]
"The practical takeaway is clear," Chen said. "Price alone should not be used to infer which model is actually cheaper."
- [16]
The researchers published their per-run cost data and code so companies could repeat the comparison on their own workloads.
- [17]
Only 11% of 396 enterprises surveyed in April and May 2026 could forecast AI costs within 10%, down from 15% in 2025, according to a report released July 29 by Benchmarkit and Mavvrik.
- [18]
Thinking tokens are the hidden reasoning a model writes before answering, and their volume helped explain the price reversals.
- [19]
Mavvrik sells AI cost-management software, and the survey figures are self-reported.
- [20]
AI work is billed by the token, and a model's rate card is a poor guide to its bill because token consumption varies between models and even between runs of the same model.
- [21]
The 336 comparisons equal every pairing of the eight models on each of the 12 tasks.
- [22]
In 230 of the 336 comparisons, the lower-listed model did not cost more.
- [23]
If the 80% discount applied evenly across input and output, Gemini 3 Flash billed about 6.9 times GPT-5.4's token volume to end up 38% more expensive.
- [24]
Gemini 3 Flash took roughly 12 times as many steps as Gemini 3.1 Pro on the failed prompt, at 14 times the cost.
- [25]
Naga's demonstration spend was about $600 an hour.
Sources
1 independent publisher whose own reporting we read for this story.
- implicator.aiCheaper AI Models Cost More in 32% of Tests, Study Finds
1 article · October 5, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.