Invest1 publisher3 min readPublished
The cheap-token trade is closing: DeepSeek's 12x price rise resets everyone's AI cost model
Cache-hit pricing went from 0.025 to 0.3 yuan, peak and off-peak rates arrived, and Moonshot, ByteDance and Alibaba are moving the same direction. Reprice now.
The Investor · Invest desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Chinese AI developer DeepSeek announced a plan to raise application programming interface (API) fees alongside the official release of "DeepSeek V4 Pro."
- DeepSeek, which drew in users with an ultra-low-price strategy, has raised its model usage fees by up to 12-fold.
- The price for "cache hits," which reuse previously processed input values, rose 12-fold, from 0.025 yuan to 0.3 yuan.
- DeepSeek introduced a differentiated pricing scheme separating peak and off-peak hours.
- DeepSeek rapidly expanded its user base with an ultra-low-price strategy set at about one-tenth the level of competitors.
Compiled by The InvestorSomething wrong?How this is made
Why it matters
DeepSeek has raised its model usage fees by up to 12-fold, announced alongside the official release of DeepSeek V4 Pro [1][2]. If your unit economics assumed Chinese inference would stay at roughly a tenth of what competitors charge [5], the assumption has an expiry date on it, because Moonshot AI and ByteDance have introduced paid pricing plans and Alibaba has signalled it will follow [8].
Look at where the 12x actually lands. The price for cache hits, which reuse previously processed input values, went from 0.025 yuan to 0.3 yuan [3]. That is exactly twelve times, an increase of 1,100 percent on the line item [10][11]. Cache hits are what reward repetitive production traffic: long system prompts, retrieval contexts, agent loops that resend the same prefix on every turn. The increase therefore falls heaviest on precisely the workloads that looked cheapest to run at scale, which is the opposite of how buyers usually model a price change.
DeepSeek also introduced differentiated pricing separating peak and off-peak hours [4]. Cost is now a function of when you call, not just how much you call. Batch enrichment, overnight evaluation runs and document backfills can be moved. Interactive, latency-bound traffic cannot, so the effective increase for a consumer-facing product is worse than the headline suggests.
One piece of arithmetic worth doing before the next board meeting. The one-tenth figure describes DeepSeek's general positioning against competitors rather than the cache line specifically [5], but twelve times one-tenth is 1.2 [12]. On that line, the arbitrage has not narrowed; it has inverted. Anyone who justified a Chinese-model dependency purely on price, and accepted the data-residency, procurement and continuity questions that came with it, is now paying for those questions rather than being paid to accept them.
The motive is not hidden. Analysts cited by Seoul Economic Daily's English edition say DeepSeek overhauled its strategy ahead of a 50 billion yuan funding round, about 10 trillion won, and an initial public offering [7]. The land-grab worked first: DeepSeek took the top spot in global token throughput from July 27 to August 2 [6]. The publication frames the whole sector as shifting its competitive axis from cost-effectiveness to securing profitability [9]. That is the ordinary sequence, and the operators who read subsidy as structural cost advantage have been here before with ride-hailing and cloud credits.
Meanwhile the hardware side of the same bill is tightening. SK Group chairman Chey Tae-won has assessed that the global memory market will face its worst-ever supply imbalance next year, and warned of "chipflation," in which rising chip prices spill over into higher prices for finished products [13][14]. He told CNBC he intends to build a front-end wafer-processing memory facility in the United States, while noting that finding a suitable site is difficult [15]. Token prices and the silicon underneath them are moving the same way at the same time.
What to watch: whether Alibaba's paid tier arrives at a level that anchors the market or merely tracks DeepSeek [8]; how deep the off-peak discount runs, since that determines whether re-architecting for scheduled batches actually pays [4]; and whether further pricing moves land before or after DeepSeek's funding round closes [7].