Leadership1 publisher3 min readPublished
Flagship output tokens now run about what GPT-3.5-level capability cost in 2022
Kion's chief executive argues in Forbes that tokens need cloud-style FinOps discipline, citing a 286-fold fall in per-token inference cost alongside LLM spending that tripled in 2025 and 93 percent of surveyed firms overrunning their AI budgets.
The Board Room · Leadership desk

What happened
- Enterprise spending on large language models tripled in 2025, according to a Forbes Tech Council column by Brian Wilson, chief executive of the cost-management vendor Kion.
- The column quotes a McKinsey survey in which 93 percent of respondents reported exceeding their AI budgets.
- About 20% of organizations curtailed their AI use because of escalating costs, particularly as their AI projects scaled upward.
Compiled by The Board RoomSomething wrong?How this is made
Why it matters
- constraint A daily spending ceiling controls the bill by rationing the work: the engineer who hits it stops until tomorrow, so a finance control shows up in delivery schedules.
- decision Executives asking what AI returned currently get an aggregate invoice, so the choice this quarter is whether to fund per-team allocation now or keep answering the ROI question with a total.
- exposure Agents that retry queries on their own move the overrun to whoever delegated model access, and naming the consumer takes per-team tracking that has to exist first.
- contradiction The consumption-sprawl diagnosis comes from a vendor of the cure, and the price and spending figures cover different years, so a buyer testing the claim has to do it on its own billing data.
The price that collapsed is the price of a fixed capability. Stanford HAI's "AI Index 2025" tracks GPT-3.5-level output, where $20 per million tokens in 2022 became 7 cents in 2024, a fall of about 286 times [1][15]. Enterprises in 2025 have moved past GPT-3.5-level output. Wilson's own range puts input for flagship models at $5 per million or more, with output tokens at three to five times input [5][6]. That is $15 to $25 per million output tokens [16]. The top of that range sits above what a million tokens of 2022-grade capability cost [17].
A tripling of spend has two candidate explanations, and the published figures leave both open. One is volume: more teams, more calls, and retries on top of both. The other is mix: the same team moving from a cheap model to a flagship one, or to a reasoning model that emits far more output tokens per answer. The price series ends in 2024 and the spending figure is for 2025 [1][2], so the two numbers cannot be multiplied into a consumption estimate.
On who is spending, the column is more concrete. AI is available to sales, marketing, finance, product, operations, engineering and executive teams rather than a relative few, and Wilson writes that costs are growing faster than teams can track [8]. "I'm seeing organizations struggle more than ever to track and control token spend," he wrote [9]. Agents make it worse by initiating activities on their own and retrying queries without assistance [11].
His remedy is the cloud playbook, applied earlier. Wilson wrote that "FinOps for AI must be an all-of-enterprise, culture-changing program, with visibility as the first step" [19]. In practice that means aggregating AI use and spending into a single view so teams can allocate at a granular level and answer the ROI question executives are asking [20], and replacing weekly reports with real-time tracking and alerting [10]. He also describes an organization that sandboxed an AI tool by giving engineers a daily spending budget, and says it worked [13].
Wilson runs the company that sells the cure. He is chief executive of Kion, whose FinOps+ platform sells governance and cost management to regulated enterprises and agencies [7]. The diagnosis can be true anyway, and a buyer can check it against its own invoices before buying anything. The 93 percent overrun figure is attributed to a McKinsey survey; the column does not name a source for the finding that about 20% of organizations curtailed AI use on cost [18].
The trade-off sits between the two remedies he offers. A daily or monthly ceiling is enforceable next week, needs no tagging scheme, and buys a stopping point and nothing else. Granular allocation takes a quarter of plumbing and is the only one of the two that answers the executive question, since the organizations that have not curtailed AI are, by his account, struggling to justify the spend with hard metrics [21]. A cap set before the allocation exists hands the tokens to whoever is loudest about the ceiling; return does not enter into it. The 20% who curtailed AI because of escalating costs hit that ceiling at enterprise scale [4].
What to watch
- Whether the FinOps Foundation's Tokenomics work produces an allocation standard vendors implement, or stops at a shared vocabulary.
- Whether the McKinsey overrun figure falls in the next survey round once 2025 budgets were reset upward.
- Whether the share of organizations curtailing AI on cost moves past 20% as agentic workloads scale.