Published Leadership3 min read
Uber's 5,000 engineers burned a year of AI tokens in four months
The binding constraint on enterprise AI is turning out to be the budget line, not the model. A pilot that works for 50 people tells you almost nothing about what happens at 5,000.
Context for builders, not their beat.See today for builders

What happened
- Uber rolled out AI coding assistants to its 5,000-strong engineering team and burned through its entire annual token allocation in just four months.
- Consuming a full annual token allocation in four months implies a consumption run rate of about three times the annual budget.
- Eight months of the year remained after the annual token allocation was exhausted.
- The cost of scaling AI initiatives does not always increase in a straight line and can often be exponential.
- A system that works brilliantly for 50 people can become expensive, risky and difficult to control when it reaches 5,000 users.
Compiled by The Board RoomSomething wrong?How this is made
Why it matters
Uber rolled out AI coding assistants to its 5,000-strong engineering organisation and exhausted its entire annual token allocation in four months, according to Bernard Marr writing in Forbes [1]. The figure that belongs on the whiteboard is not the four months but the implied run rate: roughly three times the annual budget, with eight months of the year left to fund [2] [3].
Nothing in that story is a model quality failure. The tools were used, which is what a successful rollout looks like. The failure is in the forecast. Marr's argument is that scaling costs do not rise in a straight line and can be closer to exponential [4], which means a pilot's cost per user is not a coefficient you can multiply. The illustration he uses is a system that works well for 50 people and becomes expensive, risky and hard to control at 5,000 [5], a hundredfold jump in population [6] against a cost curve nobody has actually measured at that end.
Agentic architectures make the arithmetic worse rather than better. Because agents are always on and act autonomously, Marr writes, they consume tokens far faster than non-agentic AI [7]. Firms moving from chat-style assistants to agent fleets are therefore changing the shape of the curve at the same moment they are increasing the number of people on it.
The second-order costs are also underpriced. Governance and guardrails are more onerous at scale than in a pilot, where exposure is contained to a vetted, trained group [8]. Shadow AI, meaning staff using unapproved tools in breach of policy, has already produced cybersecurity incidents serious enough to trigger regulatory action [9]. Accountability shifts too: in a pilot the buck stops with whoever is running it, while at scale customers, regulators and courts come looking for a responsible party, and a model cannot be one [10]. Regulators are increasingly treating what your AI says as a statement by your company, and boilerplate warnings that AI may make mistakes are not a defence [11]. Marr's remedy is unglamorous and cheap by comparison: document who owns output and oversight, and log every automated decision so it can be traced [12].
There is also a selection problem upstream. Pilots are frequently chosen because they demonstrate well, or because they solve a problem that is well understood but not business-critical [13]. They also attract the enthusiasts, so the cultural effects only surface when everyone is enrolled [14]. Marr cites a recent Gallup report suggesting that employees disengaged from or disgruntled about AI can themselves constitute a security risk [15].
What to watch inside your own numbers: tokens per active user per month, not total spend, because total spend hides whether the increase is more users or heavier users. Watch whether the allocation is pooled or per seat, since a pooled budget consumed by early adopters is a rationing decision made by accident. Watch the ratio of agent-initiated to human-initiated calls in any autonomous pilot before it is approved for production. And watch for a named owner attached to AI output [12]; if the pilot cannot produce one, the scaled version will not either.
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
Uber rolled out AI coding assistants to its 5,000-strong engineering team and burned through its entire annual token allocation in just four months.
- [4]
The cost of scaling AI initiatives does not always increase in a straight line and can often be exponential.
- [5]
A system that works brilliantly for 50 people can become expensive, risky and difficult to control when it reaches 5,000 users.
- [7]
Because of their always-on, autonomous nature, AI agents often burn through tokens far more quickly than non-agentic AI.
- [8]
Governance and guardrailing are far more onerous at scale than during a pilot, because pilots are self-contained with exposure limited to a vetted, trained group.
- [9]
Shadow AI, meaning workers using unapproved, unassessed tools in breach of company policies, has already caused cybersecurity incidents serious enough to trigger regulatory action.
Sources & coverage · 1 publisher
The reporting this story was synthesized from, earliest first. Every link goes to the original.
- forbes.comBernard Marr, ContributorAug 13The 5 AI Scaling Mistakes That Could Derail Your Business
Additional citations
- Bernard Marr, writing in Forbes
- Bernard Marr, Forbes
- Gallup, cited by Bernard Marr in Forbes


