Leadership1 publisher3 min readPublished
Re-baseline AI procurement on cost per completed task, not dollars per million tokens
Per-token prices fell about 200x since GPT-4's launch while US enterprise AI spend tripled to $37 billion. Tokenizer variance, reasoning tokens and tier discounts are where the bill diverges.
The Board Room · Leadership desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Menlo Ventures' 2025 State of Generative AI in the Enterprise (about 500 US enterprise decision-makers, December 2025) puts US enterprise generative AI spend at $11.5 billion in 2024, rising to $37 billion in 2025, a 3.2x year-over-year jump described as the fastest enterprise category expansion in history.
- Frontier-grade tokens are 99.5% cheaper than GPT-4's launch price; per-token, frontier-class capability is nearly two orders of magnitude cheaper than three years ago.
- When OpenAI launched GPT-4 in March 2023 the published price was $30 per million input tokens and $60 per million output, which blends to $37.50 per million at a 3:1 input-to-output ratio.
- By April 2026 DeepSeek V4-Flash was priced at $0.14 per million input tokens and $0.28 output, $0.175 blended on the same 3:1 formula, which rounds to $0.18.
- The GPT-4 to DeepSeek V4-Flash blended price ratio is 214x exact, or 208x with the V4-Flash blended price rounded, on this specific basket.
Compiled by The Board RoomSomething wrong?How this is made
Why it matters
Menlo Ventures, surveying roughly 500 US enterprise decision-makers in December 2025, put US enterprise generative AI spending at $11.5 billion in 2024 and $37 billion in 2025, a 3.2x jump it calls the fastest enterprise category expansion in history [1]. Over broadly the same period the per-token price of frontier-class capability fell by nearly two orders of magnitude [2], which is the arithmetic tell that the rate card in your procurement deck has stopped predicting the invoice.
The deflation itself is not in dispute. GPT-4 launched in March 2023 at $30 per million input tokens and $60 per million output, or $37.50 blended at a 3:1 input-to-output ratio [3]. By April 2026 DeepSeek V4-Flash was at $0.14 input and $0.28 output, $0.175 blended on the same formula [4], a ratio of 214x exact or 208x with rounding on that specific basket [5] and a 99.5% cut in the blended price [1]. Stanford's 2025 AI Index measured a 280x decline in inference cost between November 2022 and October 2024 using an equal-capability method pinned to GPT-3.5-level MMLU performance [6], and Epoch AI puts the broader trajectory at 5x to 10x a year, up to 32x for high-performance models [7].
Volume does not close the gap. For spending to triple on those unit prices, enterprises would have to be running something like 100x the workloads they ran in 2023, and Hexaware's reading of production-deployment data says they are not [13]. Three things sit between the quote and the invoice, and they compound [13]: tokenizer differences that swing the token count for the same text by roughly 0.65x to 1.35x [9], reasoning tokens that add 1.18x to 9x to the output count [10], and discounts that only apply to the workload shapes they were designed for [11]. At their extremes that is a spread of about 16x between the most and least favourable combination, before anyone negotiates [3]. The divergence is widest, per the same analysis, on reasoning-heavy workloads and Pro/Fast service tiers [12], precisely where the premium for closed models is hardest to defend.
o1 is the case in point. OpenAI shipped o1-preview in September 2024 and production o1 that December, with an internal chain of reasoning billed to the customer at the $60 per million output rate and never shown to them [15][16]. On a fixed seven-benchmark suite, Artificial Analysis measured o1 emitting 44 million tokens against GPT-4o's 5.5 million on the same problems, an 8x multiplier that took a $109 evaluation to $2,767 while the visible answer stayed the same [17]. A 25x bill on an 8x token count implies an effective per-token rate roughly 3.2x higher [4]. Those figures come from suites built to stress reasoning and will not match a production traffic mix [18].
The asymmetry matters for model selection. Because token inflation multiplies the base rate, it punishes expensive endpoints hardest [19]: a 3.5x thinking multiplier on Claude Opus 4.7's $25 output rate is an effective $87.50 per million [20], while a 1.62x multiplier on DeepSeek V4-Pro's $3.48 comes to about $5.64 [21][5], a gap of roughly 15x on the same reasoning behaviour [5].
What to watch is whether your own finance function can produce a cost per completed task rather than a cost per million tokens [8]. Until it can, reasoning tokens per task, tokenizer behaviour on your actual corpus, cache misses and service-tier routing remain unpriced variables in a contract you have already signed [14].