Leadership1 distinct publisher3 min readUpdated
Per-token prices fell about 200x since GPT-4's launch while US enterprise AI spend tripled to $37 billion. Tokenizer variance, reasoning tokens and tier discounts are where the bill diverges.
The Board Room · Leadership desk

Compiled by The Board RoomSomething wrong?How this is made
Menlo Ventures, surveying roughly 500 US enterprise decision-makers in December 2025, put US enterprise generative AI spending at $11.5 billion in 2024 and $37 billion in 2025, a 3.2x jump it calls the fastest enterprise category expansion in history [1]. Over broadly the same period the per-token price of frontier-class capability fell by nearly two orders of magnitude [2], which is the arithmetic tell that the rate card in your procurement deck has stopped predicting the invoice.
The deflation itself is not in dispute. GPT-4 launched in March 2023 at $30 per million input tokens and $60 per million output, or $37.50 blended at a 3:1 input-to-output ratio [3]. By April 2026 DeepSeek V4-Flash was at $0.14 input and $0.28 output, $0.175 blended on the same formula [4], a ratio of 214x exact or 208x with rounding on that specific basket [5] and a 99.5% cut in the blended price [1]. Stanford's 2025 AI Index measured a 280x decline in inference cost between November 2022 and October 2024 using an equal-capability method pinned to GPT-3.5-level MMLU performance [6], and Epoch AI puts the broader trajectory at 5x to 10x a year, up to 32x for high-performance models [7].
Volume does not close the gap. For spending to triple on those unit prices, enterprises would have to be running something like 100x the workloads they ran in 2023, and Hexaware's reading of production-deployment data says they are not [13]. Three things sit between the quote and the invoice, and they compound [13]: tokenizer differences that swing the token count for the same text by roughly 0.65x to 1.35x [9], reasoning tokens that add 1.18x to 9x to the output count [10], and discounts that only apply to the workload shapes they were designed for [11]. At their extremes that is a spread of about 16x between the most and least favourable combination, before anyone negotiates [3]. The divergence is widest, per the same analysis, on reasoning-heavy workloads and Pro/Fast service tiers [12], precisely where the premium for closed models is hardest to defend.
o1 is the case in point. OpenAI shipped o1-preview in September 2024 and production o1 that December, with an internal chain of reasoning billed to the customer at the $60 per million output rate and never shown to them [15][16]. On a fixed seven-benchmark suite, Artificial Analysis measured o1 emitting 44 million tokens against GPT-4o's 5.5 million on the same problems, an 8x multiplier that took a $109 evaluation to $2,767 while the visible answer stayed the same [17]. A 25x bill on an 8x token count implies an effective per-token rate roughly 3.2x higher [4]. Those figures come from suites built to stress reasoning and will not match a production traffic mix [18].
The asymmetry matters for model selection. Because token inflation multiplies the base rate, it punishes expensive endpoints hardest [19]: a 3.5x thinking multiplier on Claude Opus 4.7's $25 output rate is an effective $87.50 per million [20], while a 1.62x multiplier on DeepSeek V4-Pro's $3.48 comes to about $5.64 [21][5], a gap of roughly 15x on the same reasoning behaviour [5].
What to watch is whether your own finance function can produce a cost per completed task rather than a cost per million tokens [8]. Until it can, reasoning tokens per task, tokenizer behaviour on your actual corpus, cache misses and service-tier routing remain unpriced variables in a contract you have already signed [14].
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Frontier-grade tokens are 99.5% cheaper than GPT-4's launch price; per-token, frontier-class capability is nearly two orders of magnitude cheaper than three years ago.
When OpenAI launched GPT-4 in March 2023 the published price was $30 per million input tokens and $60 per million output, which blends to $37.50 per million at a 3:1 input-to-output ratio.
By April 2026 DeepSeek V4-Flash was priced at $0.14 per million input tokens and $0.28 output, $0.175 blended on the same 3:1 formula, which rounds to $0.18.
The GPT-4 to DeepSeek V4-Flash blended price ratio is 214x exact, or 208x with the V4-Flash blended price rounded, on this specific basket.
Independent measurement by Artificial Analysis on a fixed seven-benchmark evaluation suite found o1 produced 44 million tokens to GPT-4o's 5.5 million on the same problems, an 8x multiplier that turned a $109 evaluation into a $2,767 one; the visible portion of the answer was unchanged and the bill was 25x larger.
Token inflation is a multiplier on the existing per-token base rate, so it hurts expensive APIs more than cheap ones.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Named third-party datasets, one publisher, unverifiable rate cards
The argument is unusually well-sourced for a vendor blog: Menlo Ventures spend figures with sample size and date, Stanford AI Index and Epoch AI cost-decline series with distinct methodologies, an Artificial Analysis token measurement with absolute dollar amounts, and explicit caveats that ratios are basket-specific and multipliers benchmark-derived. It is capped by structure rather than content: everything reaches the reader through a single self-published source, and the most dramatic effective-rate examples rest on forward-dated rate cards (DeepSeek V4-Flash/V4-Pro, Claude Opus 4.7) that nothing else in the cluster corroborates.
Market conditions documented; the prescribed procurement unit is not
The underlying conditions are well attested: a $37B US enterprise spend base, reasoning models in production since December 2024 with billed hidden tokens, and rate cards moving by orders of magnitude. What is not evidenced is adoption of the thing the story actually recommends - no enterprise, vendor or benchmark in the cluster is shown re-baselining procurement on cost per completed task, and no production traffic-mix data is offered. Adoption of the problem is high; adoption of the remedy is unobserved.
Direction hedged, magnitudes stretched
Mildly overstated rather than inflated. The publisher hedges its own headline well - it labels the 208-214x ratio basket-specific, cites two other methodologies, and warns that multipliers come from reasoning-stress benchmarks. The gap comes from three places: the pivotal claim that workload volumes did not rise ~100x is asserted with no data; the most attention-grabbing ratios (up to 360x effective, $5,400/M) are built by stacking maximum thinking multipliers onto forward-dated single-source rate cards; and the 25x bill headline blends token inflation with a higher posted output rate, which the article's own arithmetic implies is roughly 3.2x of the effect.
Self-published vendor blog with adjacent advisory interest, no disclosure
The single source is a blog post on an IT-services company's own domain, framed in marketing language ('explore why AI model pricing comparison is breaking old cost benchmarks, and how dynamic usage-based pricing is redefining budgeting and transparency') with share widgets. Its conclusion - that procurement units are broken and buyers need task-level cost instrumentation and closer scrutiny of closed-source premiums - maps directly onto demand for the kind of advisory and engineering work such a publisher sells. No commercial-interest disclosure, no vendor right of reply, and no external editorial check appears in the supplied material. The score is not higher because the analysis leans on named external datasets and states caveats against its own headline.
Coherent single-source analysis, no corroboration
Confidence is limited chiefly by cluster structure: one publisher, no cross-checking, and several load-bearing prices dated in the future relative to the events described and attributable to no cited rate card. Against that, the internal arithmetic is reproducible (the 99.5%/214x, $87.50 and ~$5.64 figures all check out), the external datasets are named and dated, and the publisher flags its own methodological limits. The mechanisms described are therefore more trustworthy than the specific magnitudes.
product
A 27B laptop model scores like a rented one, and thinks three times as hard to do it1 distinct publisher
build
Your Multi-Key Failover Is The Most Expensive Line On Your Coding Agent Bill1 distinct publisher
invest
Google Ships Flash Instead of Pro While OpenAI Loses Its Two Best Operators1 distinct publisher
product
A Beijing bar gives away DeepSeek tokens. Your per-token price list is the collateral damage.1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 18, 2026