Skip to content

Invest1 publisher3 min readPublished

DeepSeek's 12x cached-token rise ends the cheap-endpoint era for Chinese inference

V4 Pro shipped with an API price increase of up to 12-fold, and Moonshot AI and ByteDance have added paid tiers. Anyone whose margins assumed cheap Chinese endpoints needs to re-model.

The Investor · Invest desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • Chinese AI developer DeepSeek unveiled the official version of DeepSeek V4 Pro and announced an API price increase at the same time.
  • DeepSeek raised its application programming interface (API) fees by up to 12-fold alongside the launch of DeepSeek V4 Pro.
  • During peak hours (2 a.m. to 9 a.m. and 2 p.m. to 6 p.m.), input costs 9 yuan and output 27 yuan per million tokens.
  • Off-peak hours are priced at half the peak rate.
  • The cache-hit price for reusing previously output input values jumped 12-fold, from 0.025 yuan to 0.3 yuan.

Compiled by The InvestorSomething wrong?How this is made

Why it matters

DeepSeek released the official version of DeepSeek V4 Pro and, at the same time, announced an API price increase of up to 12-fold [1][2]. The lab that took global token-throughput leadership by pricing at roughly one-tenth of competitors is no longer subsidising volume, and its domestic peers are moving in the same direction [6][8].

The new sheet: during peak hours, defined as 2 a.m. to 9 a.m. and 2 p.m. to 6 p.m., input costs 9 yuan per million tokens and output 27 yuan, with off-peak hours priced at half that [3][4]. Off-peak therefore lands at 4.5 yuan input and 13.5 yuan output [1], and a symmetric job of one million tokens in and one million out costs 36 yuan at peak against 18 off-peak [4]. Output is priced at three times input [2], so verbose agent chains are penalised more than long prompts.

The headline 12x is not the per-token list rate. It is the cache-hit price for reused input, which went from 0.025 yuan to 0.3 yuan per million tokens [5]. That is the sharpest increase in the announcement [5], and it falls on exactly the architectures that were engineered around it: long fixed system prompts, retrieval pipelines that resend the same corpus, and multi-turn agents that replay context on every step. If your cost model assumed cached reads were effectively free, the input side of your bill is now roughly a thirtieth of the peak uncached rate rather than a rounding error [5]. Re-run the arithmetic before quoting a customer again.

There is one genuine lever left. The peak windows cover 11 of 24 hours [3], so batch and offline work that can be scheduled into the other 13 pays half [4][3]. That is a scheduling problem, not a modelling one, and it is cheaper to solve than a migration.

On motive, the reporting is analyst inference rather than company statement: sedaily.com reports that analysts attribute the strategy reversal to a 50 billion yuan funding round, about 10 trillion won, and a planned initial public offering [7]. The same piece notes Moonshot AI and ByteDance introducing paid pricing plans and Alibaba weighing a similar move, with the competitive axis among Chinese AI firms shifting from cost efficiency to profitability [8][9]. Pre-IPO, a book of below-cost inference is a liability; margin is the number that gets underwritten.

The counter-argument from U.S. vendors, per the same report, is cost-per-outcome: the same task completed with fewer tokens can cost less in total even at a higher unit price [10]. That argument was easy to dismiss at a tenth of the price. At the new cache-hit rate it deserves an actual bake-off on your own workload, measured in completed tasks rather than tokens purchased.

What to watch: whether Alibaba follows Moonshot and ByteDance into paid tiers [8], which would remove the last obvious cheap substitute; whether the off-peak half-price band survives the next revision, since it is the easiest concession to withdraw [4]; and whether DeepSeek holds throughput leadership after the increase, given that the ranking it won from the 27th of last month through the 2nd of this month was earned on the old price [6]. Volume that was bought with price tends to leave with it.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories