Leadership1 distinct publisher3 min readPublished
Metered spend broke the per-seat budget at Uber, and a dollar ceiling only rations the invoice. Commonwealth Bank shows the version of the same argument that survives a CFO's questions.
The Board Room · Leadership desk

Compiled by The Board RoomSomething wrong?How this is made
The two-hour demo is the detail worth sitting with. Twelve hundred dollars of tokens went out the door in a single session [4], which is about 80% of what one engineer is now permitted for an entire month [2]. Nothing in that number was visible to anyone until the invoice arrived. The same arithmetic explains the headline: consuming a twelve-month budget in four months is a run rate three times the plan [5], and it took a quarter of the year to notice.
That is the governance failure, and it is separate from the value question. Deloitte's 2026 enterprise survey puts 74% of organisations hoping AI will grow revenue against 20% doing it today, a gap of 54 points [4]. Meanwhile more than a third are already banking efficiency gains [9]. Efficiency is the return most finance teams can actually document, and it is also the return a metered bill eats first, because savings and token spend land on the same side of the ledger.
Commonwealth Bank's roughly $2.4bn a year on technology and capability [10] is, on its own, a number that settles no argument. What makes it defensible is the line items underneath it: a generative tool called Compass that has handled over 500,000 banker queries and lets frontline staff answer complex questions about three times faster [12]. Chief financial officer Alan Docherty's stated logic is that productivity and better customer outcomes feed top-line revenue, which is what justifies the next round of investment [13]. That is a sentence a CFO can repeat in February and be held to in August.
There is a second test that has nothing to do with dollars. Harrison.ai trained its radiology models on more than a million clinical studies and 550 million expert annotations from over 140 radiologists [14], and independent testing has shown its chest X-ray tool lifting diagnostic accuracy by 45% [15], with more than half of Australia's radiologists now having access [16]. The asset there is the annotation corpus, not the model. Rob Versaw's argument is that an organisation whose AI reasons from the open web instead of its own context pays premium prices for generic output a rival can buy just as cheaply [19]. Worth noting where the argument comes from: Versaw works in product strategy at Dynatrace [18], an observability vendor, so per-workflow instrumentation of spend is a conclusion that suits his employer. It does not make the Uber figures less real.
The gap in all of this material is a denominator. There are caps, totals, monthly ceilings and query counts, and not one cost per merged change, resolved ticket or shipped feature [6]. Until an engineering organisation publishes that rate, "per-token value" stays an argument about which workflows get named, not a measurement anyone can check.
Ranked by verification strength, evidence, and original report placement.
In April, Uber's chief technology officer told The Information (reported via Forbes) that the company had spent its entire 2026 budget for AI coding tools in four months.
Roughly 5,000 engineers were using Uber's AI coding tool program.
Some Uber "power users" were running up to $2,000 a month in tokens.
By June, Uber had capped individual AI coding tool spending at $1,500 a month.
Uber's chief operating officer conceded it was hard to draw a line between token spend and features customers actually notice.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Single vendor-authored op-ed, key figures secondhand
Everything rests on one Forbes Technology Council contribution written by a Dynatrace product strategy executive. The Uber numbers are relayed secondhand from The Information; the CBA, Harrison.ai and Splose figures are company-disclosed and not independently checked here; the 5x-30x agentic token multiple and the 45% accuracy lift carry no named study. The disclosures themselves are specific and internally consistent, which keeps the score from being lower, but nothing in the cluster is corroborated by a second publisher.
Real deployments at scale, thin outcome measurement
The disclosed footprint is substantial and named: roughly 5,000 Uber engineers on paid AI coding tools with spend large enough to exhaust an annual budget, CBA running AI against more than 20 million payments a day plus 500,000 Compass queries, majority access to Harrison.ai's chest X-ray tool among Australian radiologists, and clinicians paying for Splose's AI tier. Adoption is clearly past pilot stage. It is scored below the top band because usage counts are not paired with per-unit output measures and the Uber cap suggests consumption was being restrained rather than optimised.
Sober thesis, unverified proof points
The article's argument runs against AI hype - it insists on attribution, ownership and measured outcomes - and the Deloitte 54-point expectation gap and the Uber COO's admission are deflationary. The overstatement sits in the supporting exhibits: a 45% accuracy lift from unnamed testing, a three-times-faster claim, a 20% fraud-loss reduction credited to AI without a counterfactual, and a 5x-30x token multiple with no source. The prescription (name the workflow, see the cost early) is also asserted to fix the economics without any evidence that it has. Mildly positive, not severe.
Vendor-authored council post prescribing its own category
The byline discloses that the author works in product strategy at Dynatrace, an observability vendor, and the piece is published through Forbes' invitation-only Technology Council rather than as reported journalism. Its central prescription - make workflow-level cost and value visible before the invoice arrives, and prefer proprietary context over general-web output - maps directly onto the commercial category the author's employer sells. The disclosure is explicit and the factual content is checkable, which caps the score below the extreme.
Specific disclosures, no corroboration
Confidence is moderate: the individual figures are precise, dated and attributable to named executives or results announcements, which makes them usable as directional signal. But the cluster has one publisher, an incentive-aligned author, secondhand sourcing on the anchor facts, and no unit-cost or independent-verification layer, so the central conclusion - that a spend cap is not the fix - remains an argument rather than a demonstrated finding.
product
Canva's forecast cut turns model routing into a product line item1 distinct publisher
leadership
EY Answers The AI-ROI Question With An Org Chart: One Office, One Budget1 distinct publisher
invest
A CEO's $1,000 weekend, and the auto-renew setting that made it possible1 distinct publisher
product
Anthropic's Ode has bought two consultancies in four months. That is the product.1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
forbes.com
1 article · August 26, 2026