Invest1 distinct publisher2 min readPublished
OpenRouter says DeepSeek went from 9% to 18% of tokens in six months, with agent loops doing most of the work. Token share and spend share are not the same league table.
The Investor · Invest desk
Compiled by The InvestorSomething wrong?How this is made
Agentic requests burn about 15 times the tokens of a human request, according to OpenRouter's own data [12], and that multiplier is the mechanism behind everything else. On a blended million in and million out, V4 Flash's cheapest endpoint runs about $0.27 against roughly $35 for GPT-5.5, a gap of about 130x [2]; on output alone it is closer to 167x [1]. The ratio does not change when a workload turns agentic, but the absolute money does, because the same gap is applied to fifteen times the tokens [5]. Anything that runs in a loop and calls tools stops being a model choice and becomes a procurement decision about price per token.
Inside DeepSeek's own traffic the two populations behaved differently, which is the more useful evidence. Human sessions stayed on V3.2 through the first five months of the year [14], so this is not a broad swing in taste; it is agent loops shopping on cost. Hobbyists now route close to a third of their tokens to DeepSeek models [7], and OpenRouter says users at AI-native companies and large organisations also sent considerably more traffic there in early June than they did in January [8].
Two things temper the read. The volume jump did not produce a matching jump in share of spend [10], so the table DeepSeek now tops is the one denominated in the cheapest unit available. And the segmentation carrying the agentic argument is OpenRouter's own: a seven-signal weighted composite scored at the API key level, with tool call rate, turn count and gap timing among the inputs [11]. That is a reasonable proxy rather than an observation of what those keys were doing, and the party publishing it sells routing. The same post puts DeepSeek at nearly 20% at the start of June [5] and at 18% in its January-to-June comparison [4], so treat the figure as a band.
The country arithmetic is the number worth writing down. If American models held about three quarters of tokens through 2025 [16] and Chinese models passed them in early June [17], Chinese share moved from no more than roughly a quarter to above half: a swing of at least 25 percentage points in under eighteen months, on a base that was itself growing [4]. The named Chinese gainers alongside DeepSeek were Xiaomi, Minimax and Tencent, and OpenRouter attributes the offsetting losses to Google and OpenAI [6]. Read plainly, buyers substituted on price at the precise moment their workloads began consuming fifteen times more of the product.
Ranked by verification strength, evidence, and original report placement.
DeepSeek began 2026 holding just under 10% of weekly token flow across OpenRouter.
As agentic work took hold in February and March 2026, DeepSeek's share fell to 5% of total OpenRouter tokens, squeezed by proprietary models above and other open-source LLMs below.
A direct January-to-June 2026 comparison shows DeepSeek's OpenRouter token share going from 9% to 18%.
By the start of June, DeepSeek had earned nearly 20% of token share and has been the top model on OpenRouter since mid-May.
DeepSeek V4 Flash on the cheapest endpoint costs $0.09 input and $0.18 output per million tokens; GPT-5.5 is priced at $5 input and $30 output per million tokens.
The upsurge in V4 token volume has not resulted in an identical spike in share of spend.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Large first-party telemetry, zero independent corroboration
The numbers rest on a disclosed methodology and a very large sample (over 450 trillion tokens of OpenRouter request logs, Jan 1 to Jun 14, 2026), which is strong for the narrow question of what flowed through OpenRouter. But the cluster contains exactly one source, that source is the measured venue itself, no underlying data or chart values are reproducible, one headline figure is stated two ways (18% and 'nearly 20%'), and the quality assertion behind 'sufficient for agentic work' has no benchmark evidence at all.
Real routed traffic, broad user mix, one venue
This is observed production routing rather than intent: DeepSeek reached roughly 18-20% of tokens and top-author position on OpenRouter, V4-Flash took 70% of DeepSeek agentic flow within a month of release, hobbyists route nearly a third of tokens to DeepSeek, and AI-native and large-organization keys increased DeepSeek traffic. Adoption is high but bounded to one aggregator's marketplace, and volume adoption is explicitly not spend adoption.
Volume story outruns the unpublished spend story
The measured claims are mostly modest and well instrumented, and the post itself flags the volume-versus-spend divergence, which pulls the gap toward zero. It stays positive because the strongest framings are the least evidenced: the ~130x price advantage uses V4 Flash's cheapest endpoint rather than a typical routed price, 'best in class cost effectiveness to output quality' has no quality measurement behind it, the market-level 'Chinese models surpassed American ones' conclusion is drawn from one router's volume, and the forward statement that V4 'is certain to be part of that mix' is asserted rather than evidenced.
Marketplace publishing data that validates its own model
OpenRouter is simultaneously the data source, the venue measured, and a commercial beneficiary of the conclusion. Its business is multi-model routing and price arbitrage, so a narrative in which cheap open-weight models displace premium American incumbents directly promotes its value proposition; the post closes by citing a WSJ story featuring OpenRouter data about startups mixing models to avoid premium prices. No conflict statement or independent audit is offered, though the disclosed methodology and the volunteered volume-versus-spend caveat cut against pure promotion.
Trust the router-level facts, not the market conclusions
Confidence is moderate: the platform-scoped facts (release date, share trajectory shape, V4-Flash's capture of DeepSeek agentic flow, agentic token intensity, posted prices) are specific, internally coherent and drawn from a huge sample, so they are likely accurate for OpenRouter. Confidence is held down by the single-source cluster with a conflicted publisher, the unreconciled 18%/'nearly 20%' figures, the missing spend-share quantification, and the extrapolation from one marketplace to US-versus-China model share overall.
build
1.5% of Hugging Face repos take 99.2% of downloads, and the ceiling is Chinese1 distinct publisher
invest
Chinese models now carry 60% of OpenRouter traffic, and 58% of what US firms route1 distinct publisher
build
Your Multi-Key Failover Is The Most Expensive Line On Your Coding Agent Bill1 distinct publisher
build
Hugging Face's $13B process puts most teams' model pipeline under a single owner2 distinct publishers
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 25, 2026