Leadership1 publisher3 min readPublished
Amazon ran an AI project five months before catching an 860% overrun
The figures driving 2026's AI budget panic reach operators secondhand, through a vendor's column citing Fortune and TechSpot. The one case with dates attached puts the cost in the five months before anyone checked.
The Board Room · Leadership desk
What happened
- A Forbes Tech Council column by NinjaCat cofounder Paul Deraval quotes Fortune reporting that Uber burned through its entire 2026 AI budget in four months.
- The same column places Uber among the nearly 70% of companies it says reported AI cost overruns in 2026; it does not name the survey.
- Citing TechSpot, it describes an Amazon project using Anthropic's Claude Sonnet to match author information to product listings that ran five months before the overrun was detected.
Compiled by The Board RoomSomething wrong?How this is made
Why it matters
- decision A quarter spent negotiating a lower price per token buys a discount on volume the buyer still cannot see; halving Amazon's token price would have left a bill 4.8 times plan.
- constraint Token metering pushes the control point down to the workflow, so account-level spend tracking cannot separate a busy month from a runaway job until an invoice arrives.
- exposure Anyone who swaps to a cheaper model to hit a cost target inherits its error rate, and the column credits the erased-savings finding to one study it does not identify.
- contradiction The prescription comes from a vendor chief executive writing in a contributor column, which leaves one dated case carrying the operations argument.
Take the two dated cases at face value and the sizes are not comparable. If Uber consumed a full-year budget in four months, it spent at roughly three times its planned annual rate [17]. Amazon's is steeper. A bill 860 percent above budget means the project finished at 9.6 times plan, which puts the original budget near $187,500 and the overage near $1.61m [14][15]. Across the five months the project ran, that is about $360,000 a month against a plan of about $37,500 [16].
The column's account of how it ran that long is about the unit of measurement. Usage metered by tokens is hard for enterprises to track and control, Deraval writes, and the Amazon case is his illustration [5]. His diagnostic argument is that token counts describe behaviour, not just price: where a model spends effort, which workflows demand excessive reasoning, where prompts lack clarity, where context forces computation nobody asked for [7]. He compares a high AI bill to a summer energy bill: "Closing the window is probably a better solution than turning off the air conditioner," he wrote [8].
Paul Deraval is cofounder and chief executive of NinjaCat, and the piece ran as a Forbes Tech Council column [10]. The 70 percent overrun figure arrives without a named survey, and the finding that cheap models with slightly higher error rates erase their own savings is credited only to "one study" [2][9]. On both, the record is an assertion.
The Amazon case does not need the argument around it. Five months of undetected spend on a workflow matching author information to product listings is a reporting failure at the workflow level, and no model price fixes it [3]. The evidence carries one part of Deraval's sequencing claim: the largest costs rarely come from model selection alone [6]. His related claims are reasonable but untested here, including the case for treating models as specialists rather than by seniority, and the warning against leading with the upfront cost of open-weight options such as DeepSeek or Kimi K3 [12][11].
Which sets the order of operations for a quarter that has AI cost on the agenda. Halving the price per token on the Amazon workflow would still have produced a bill 4.8 times the plan, because the overrun came from volume that went uncounted [19]. Model shopping is the cheaper project to run and the easier one to present, since it ends in a rate card. Attribution by workflow needs a threshold that stops a job, an owner, and a number that owner defends every month.
The trade-off is real, and what it costs is attention. Instrumenting token spend per workflow costs engineering time this quarter and yields nothing visible if no workflow misbehaves. Skipping it costs nothing this quarter and leaves the next five-month drift to be found by an invoice.
What to watch
- Whether Uber or Amazon confirms these figures directly; both reached the column secondhand, via Fortune and TechSpot.
- Whether a named survey with a sample size and question wording ever turns up behind the 70% overrun figure.
- Whether model and cloud vendors start attributing token spend by workflow in billing, instead of at the account level.