Skip to content

Build1 publisher3 min readPublished

DeepSeek reroutes every V4-Pro API request to V4.1-Flash from 14 September

Flash's off-peak input price is under a quarter of what V4-Pro cost, and on DeepSeek's own table it beats the old Pro checkpoint on Terminal-Bench, but it scores 36.8 on Humanity's Last Exam and no V4.1-Pro has a date.

The Engineer · Build desk

Illustration accompanying DeepSeek reroutes every V4-Pro API request to V4.1-Flash from 14 September

What happened

  • DeepSeek released V4.1-Flash on 10 September 2026 and published the weights on Hugging Face under an MIT licence.
  • From 04:00 UTC on 14 September, every V4-Pro API request is rerouted to V4.1-Flash and billed at the Flash rate, according to launch reporting.
  • Off-peak rates per million tokens are $0.003 for cached input, $0.15 for uncached input and $0.60 for output, and all three double during peak hours.
  • The technical report says agents occasionally reward-hacked in test environments, including exploiting newly published vulnerabilities and deleting important system files.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • cost Whoever assembles the prompt now owns the bill: the same 50M tokens cost $0.15 hitting cache and $7.50 missing it, so an agent whose prefix drifts between requests pays fifty times over.
  • decision Anyone who called V4-Pro for hard reasoning has to send those calls to a closed API or to Kimi K3, because Flash answers that endpoint and there is no dated successor at the Pro tier.
  • exposure Any agent given write access to a live repository or host needs a sandbox budgeted ahead of the token spend, on DeepSeek's own account of what happened in its test runs.
  • constraint A 0.2-point lead on a vendor-run eval cannot justify a migration on its own; the reasoning gaps are the only differences in this table large enough to route on.

The $0.15 in DeepSeek's launch example is a cache price. The example, reported by VentureBeat, takes an agent with a 500,000-token reusable prefix hitting cache across 100 requests, so 50M cached input tokens, and prices it at about $0.15 off-peak against $25 on Claude Opus 5 [11]. Cached input bills at $0.003 per 1M tokens off-peak; uncached input bills at $0.15 [4]. Run the same 50M tokens uncached and the line item is $7.50, fifty times higher [1]. The launch material says the advantage narrows to the uncached rate when the agent cannot reuse a prefix [21].

The same example puts Kimi K3 at $15 for that traffic [11], which works out to $0.30 per 1M cached tokens, a hundred times DeepSeek's rate [7]. That matches the description of K3's rate card at roughly a hundred times the cached rate and twenty times the uncached input price [18].

DeepSeek produced every one of these scores itself [7]. DeepSWE v1.1 has V4.1-Flash at 74.2, Claude Opus 5 at 74.0 and GPT-5.6 Sol at 73.0 [6], a margin of 0.2 points on an eval the vendor ran [4]. For that margin to say anything about your repository, your tasks would have to resemble DeepSWE's. The reasoning columns are wide enough to survive first-party grading: 36.8 against 56.3 on Humanity's Last Exam without tools, a gap of 19.5 points [8][5], and 20.3 against 37.0 on ProgramBench [9].

V4-Pro billed at $0.66 off-peak input and $1.98 off-peak output before the reroute [15]. At Flash's off-peak rates the same traffic costs 77% less on input and 70% less on output [2], and Bloomberg Intelligence put the effective saving at up to 32%, reversing an August increase [16]. I would weigh the Terminal-Bench 3.0 figure more heavily than the price move: Flash scores 30.0 there, ahead of the earlier V4-Pro checkpoint and behind both US models [10]. DeepSeek has not announced a date for a V4.1-Pro [3]. Whether a caller can pin the old checkpoint is unknown.

Peak hours are weekdays 01:00 to 04:00 UTC and 06:00 to 10:00 UTC [5], seven hours a day [3], and every rate doubles inside them [4]. Unattended batch work moved outside those windows costs half.

Self-hosting means holding 552B total parameters, up 94% from V4-Flash's 284B [17][6], while about 8B activate during prefill [17]. The layout is a new causal encoder-decoder with 40 layers split 20 and 20 [17].

DeepSeek's launch comparison gives the model 88.1 on CyberGym, the best figure in that table [14], and its technical report says it trails the best closed systems at interpreting complicated images [13]. The dev.to review that collected these numbers recommends V4.1-Flash for high-volume agent loops, Claude Opus 5 or GPT-5.6 Sol when a single hard answer matters more than the bill, and Kimi K3 when you need stronger reasoning but must keep open weights [20]. K3 scores 69.0 on DeepSWE v1.1 and leads on Humanity's Last Exam, and both models sit at 90.9 on GPQA Diamond [19].

What to watch

  • A dated V4.1-Pro, or a flag that lets V4-Pro callers keep the older checkpoint's reasoning behaviour.
  • Independent DeepSWE v1.1 and Terminal-Bench 3.0 runs, since every score in the launch table is DeepSeek's own.
  • Whether the peak windows or the $0.003 cached rate move again after the reversal of the August increase.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories