Invest1 distinct publisher3 min readPublished
Vercel's June data shows enterprise AI volume and spend pulling apart, with cheap open models absorbing routine work while Anthropic holds 61% of the money and 72% or more of the jobs that hurt when they go wrong.
The Investor · Invest desk

Compiled by The InvestorSomething wrong?How this is made
The ratio behind the 29% share is the interesting number: nearly a third of tokens for just under 4% of spend [1] works out to roughly an eighth of the average dollar-per-token on the gateway, and Vercel puts open-weight pricing at about a tenth of the platform average [8]. That is a discount deep enough that a buyer moving a workload does not need a procurement argument, only a routing rule.
Which is why the flat price per token [2] is the line worth pausing on. Cheap volume tripling its share since April [7] should have dragged the blended average down, and it would have, except closed-weight frontier prices rose about 12% per token in the same month [9] and cancelled it out. The average held still because two things moved in opposite directions, not because anything settled. Strip out the frontier increase and the same June looks like deflation; strip out the open-weight growth and it looks like a price rise. Same month, three readings.
Anthropic's position is the part I would not short. It took 61% of spend on 32% of tokens [3], down from 65% in May but in line with April [11], and 72% or more of spend in coding agents, back-office agents and app generation [4] - the work where a wrong answer costs real money. If a buyer's cheap tier is absorbing summarisation and classification, the frontier bill is not being cut, it is being protected. The pattern with OpenAI is the cleaner tell: token share fell from 12.5% to 10.3% while spend share rose from 13.3% to 16.1% [12], which Vercel reads as cost per token up about 50% relative to the market in a month [13]. Less volume, costlier work.
This is probably wrong, but I read the DeepSeek climb to 22.6% of tokens, within two points of Google [5], as less about model quality than about the absence of any switching cost worth naming, and GLM 5.2 is the evidence: MIT-licensed, priced near a fifth of Opus 4.8, 50x daily token growth from June 16 to month end, and 76% of its own family's June tokens in barely two weeks [14][15], against Gemini 3.1 Pro needing a second month to do the same [16]. Or rather, the more interesting version of the claim is that adoption is now faster than evaluation, since Vercel notes customer count grew far faster than token usage [17] - people are trying it before they trust it with anything.
This same month reads differently depending on which thread you pull. The frontier price increases could be temporary, in which case the blended price falls hard once they stop and the volume story becomes a straightforward cost win. Or only about one in eight enterprise customers runs an open-weight model in production [10], which would make the 29% a small cohort routing aggressively rather than a broad migration - one that stalls. Or policy decides it: a US export-control directive suspended access to Claude Fable 5 for the rest of June after it reached 22% of Opus 4.8 request volume in four days [6], and Chinese labs already take two-thirds of gateway video spend [18], so the cheap tier is one directive away from being unavailable rather than merely cheap.
What would prove the decoupling thesis wrong is simple enough to check: if Anthropic's high-stakes spend share slips below 72% while open-weight volume keeps climbing, then buyers are not protecting a tier, they are replacing one.
Ranked by verification strength, evidence, and original report placement.
Claude Fable 5 was released June 9 and reached 22% of Opus 4.8's request volume in four days; a US export-control directive suspended access for the rest of the month.
OpenAI's gateway token share fell from 12.5% to 10.3% while its spend share rose from 13.3% to 16.1%.
Vercel says OpenAI's cost per token rose about 50% relative to the market in one month, with customers sending less volume but costlier work.
From June 16 API availability through month end GLM 5.2's daily token volume grew about 50x, ranked #11 by tokens in the final week and as high as #7 on single days, and took 76% of its family's June tokens in barely two weeks.
Gemini 3.1 Pro, previously the fastest in-family migration Vercel had charted, took until its second month to reach a similar in-family share.
GLM 5.2's customer count grew far faster than its token usage in the two weeks after launch.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 30, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
product
Incogni ranks 13 AI assistants by privacy risk: bigger is worse, except ChatGPT1 distinct publisher
invest
Speed becomes a SKU: OpenAI and Google put a separate price on latency3 distinct publishers
build
Grok 4.6 lands in Copilot two days after launch, and the model picker becomes a procurement problem1 distinct publisher
leadership
A Government Switched Off Two Frontier Models. Your Board Will Want The Fallback Plan.1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Unusually specific, entirely self-reported
The granularity is real — month-over-month share moves per lab, per-modality splits, an in-family migration curve dated to the day — and telemetry from the router is the right instrument for these questions. But it is one instrument, held by one party, and it reports only ratios: no token or dollar totals, no customer count, no definition of a 'high-stakes use case'. The one place the arithmetic can be checked from the inside, it wobbles.
Real production traffic, one gateway wide
This is not intent data. Tokens actually billed, models actually called: open weights at 29% of volume, one in eight customers running them in production, GLM 5.2 going from API availability to a top-eleven position inside two weeks. The ceiling on the score is scope, not substance — everything observed is traffic that chose to pass through Vercel, whose customer base skews toward web application teams rather than, say, banks running models in their own data centres.
Volume framing outruns the dollars
The headline number is true and the framing leans on it hard. Twenty-nine percent of tokens for a twenty-fifth of the dollars means open weights have taken the cheap work, and Anthropic's 61% of spend and 72%-plus of the consequential jobs say the valuable work has not moved at all. Vercel's projection that an open-weight lab will 'soon' rank second by volume is a two-month trend line extended without a date, and second place by tokens at a tenth of the price is not second place by revenue. The overstatement is in emphasis, not in the figures.
The router publishes the routing data
Vercel sells the multi-model gateway, and the report's conclusion — that leading every modality now takes at least three labs, and that the disciplined move is cheap volume here, frontier risk there — is a description of what its product exists to do. That does not make the telemetry wrong; a vendor with this data has little to gain from misreporting shares its customers can audit on their own invoices. It does explain which findings got a chart and which questions, like whether the cheap models produced acceptable output, never came up.
Coherent, unchecked, and two months late
Hold this at arm's length. The internal logic hangs together and the specificity is hard to fake, but June's data reached readers at the end of August, no second party has reported an overlapping month, and the most consequential passage — the export-control directive that pulled Claude Fable 5 — is truncated before it says who ordered what. Directionally I would use it; I would not put a decimal point from it in a board deck without a second read.