Skip to content

LeadershipNot yet confirmed elsewhere1 publisher3 min readPublished

Lower-priced AI models ran up bigger bills in 32% of head-to-head tests

Researchers led by Lingjiao Chen of Stanford found the AI model with the lower list price cost more in 106 of 336 head-to-head comparisons. For anyone setting an AI budget, the finding means per-task cost has to be measured on their own workloads before a model is chosen.

The Board Room · Leadership desk

How we use AISend a correction

Illustration accompanying Lower-priced AI models ran up bigger bills in 32% of head-to-head tests
Generated illustration

What happened

  • At May 1, 2026 list prices, Gemini 3 Flash cost 80% less than GPT-5.4, but across the study's tasks it came out 38% more expensive.
  • On one MMLU-Pro problem, Gemini 3 Flash wrote more than 60,000 hidden thinking tokens, while GPT-5.4 solved the same problem with 25.
  • Running the same query again and again on one model, the costliest run cost up to 9.7 times as much as the cheapest.
  • Uber had exhausted its entire 2026 AI budget by about four months into the year.
  • Only 11% of 396 enterprises surveyed by Benchmarkit and Mavvrik in April and May could forecast AI costs within 10%, down from 15% in 2025.

Compiled by The Board RoomSomething wrong?How this is made

Why it matters

  • decision Model selection now needs a costed trial on a team's own tasks before signing, and the researchers' published per-run data and code make that trial cheaper to set up.
  • constraint Run-to-run variance survives re-prompting, so measured usage gives a cost range per query and fixed per-task chargebacks to internal teams get harder to set.
  • cost In agent workloads a cheap model that loops is billed for every turn, so the bill for a failed run can exceed a pricier model's bill for a finished one.

A rate card prices tokens. What a task costs depends on how many tokens the model decides to spend on it, and that count varies between models and between runs of the same model [20]. For Gemini 3 Flash's 80% discount to finish as a 38% premium over GPT-5.4, Flash had to bill about 6.9 times the token volume, assuming the discount held evenly across input and output [23]. The researchers traced much of that volume to thinking tokens, the hidden reasoning a model writes before it answers [18].

Agent work widens the gap. When a model works through tools or a computer environment, each turn can carry the earlier conversation back in as input, and in those tasks the number of turns drove cost [12]. On one prompt, Gemini 3.1 Pro finished in 85 steps for about $1. Gemini 3 Flash took nearly 1,000 steps, ran up $14 in token charges and failed [13]. That is roughly 12 times the steps for 14 times the money [24].

A finance lead could fairly read the headline result the other way, since the cheaper-listed model did not cost more in 230 of the 336 comparisons [22]. The 336 matches every pairing of the eight models on each of the 12 tasks [8][21], so each reversal belongs to one pair on one task. No model was consistently the cheapest or the most expensive; the ranking changed from task to task [9]. A buyer who knows the rate card was right about two times in three still cannot tell which of their own tasks sit in the other third.

Measuring on a team's own tasks helps more with choosing a model than with fixing a budget. The paper describes an irreducible noise floor: models can follow different reasoning paths even when their inputs stay fixed, and re-prompting does not remove the variation [2]. In a follow-up on two programming prompts, models from Anthropic, Google and OpenAI each got the same prompt five times and produced a different cost each time [14]. "The practical takeaway is clear," said Lingjiao Chen, a researcher at Stanford University and Microsoft Research [7]. "Price alone should not be used to infer which model is actually cheaper." [15] The team published its per-run cost data and code so companies can repeat the comparison on their own workloads [16].

Evidence that companies are missing their forecasts is of uneven quality. The survey comes from a vendor: Mavvrik sells AI cost-management software, and the figures are self-reported [19]. Uber's is a single company's account. Its chief technology officer, Praveen Neppalli Naga, used $1,200 of tokens in a two-hour demonstration of the company's coding tools in April [3], about $600 an hour [25]. "I'm back to the drawing board, because the budget I thought I would need is blown away already," Naga said [5]. The account of Uber's exhausted budget does not say how much of the overrun came from model choice and how much from volume [4].

The trade-off this quarter is between paying for a measurement run and picking from the rate card. The run costs tokens and engineering time now. Each task has to be repeated enough times to show the spread between cheap and expensive runs. I think the output worth paying for is a cost range per task type with its tail included, since that range is what next quarter's budget has to absorb. A team that skips the run sets that budget from a price comparison that reversed in 106 of 336 tests [6].

What to watch

  • Replications that run the researchers' published per-run data and code against production workloads, and whether their reversal rates land near the study's 32%.
  • Any explanation from Uber of what drove its 2026 AI budget overrun and how it reset the budget.
  • Provider moves to cap or itemise thinking tokens on bills, the volume the study found helped explain the reversals.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence62
Adoption
Insufficient
Hype gap+8
Incentives45
Confidence58
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    In the May revision, repeated runs of an identical query on the same model varied by up to 9.7 times between the cheapest and most expensive run.

  2. [2]

    The paper describes an irreducible noise floor: models can follow different reasoning paths even when inputs stay fixed, and the variation cannot be removed by re-prompting.

  3. [3]

    In April 2026, Uber chief technology officer Praveen Neppalli Naga demonstrated the company's AI coding tools and over two hours used $1,200 worth of tokens.

    ReportedSupportedSource: implicator.aiView cited source

Sources

1 independent publisher whose own reporting we read for this story.

  1. implicator.ai

    1 article · October 5, 2026

    Cheaper AI Models Cost More in 32% of Tests, Study Finds

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Loading related stories