Skip to content

BuildNot yet confirmed elsewhere1 publisher2 min readPublished Updated

DeepSeek V4.1 Flash's 36x discount to Claude Opus 5 depends on off-peak hours and low effort

DeepSeek V4.1 Flash costs $2.10 off-peak for an agent run that costs $75 on Claude Opus 5, about 36x less, according to a dev.to analysis. Paying peak-hour rates halves that gap to about 18x, and running at max effort cuts it to about 15x.

The Engineer · Build desk

How we use AISend a correction

Illustration accompanying DeepSeek V4.1 Flash's 36x discount to Claude Opus 5 depends on off-peak hours and low effort
Generated illustration
Agent run: $75 on Opus 5, $2.10 on Flash off-peak Cost of one agent run at list prices, before cache hits, per the dev.to post's math. Flash at peak-hour rates is shown alongside off-peak.

Bar chart of the cost of one agent run of 10M input and 1M output tokens: Claude Opus 5 $75, DeepSeek V4.1 Flash at peak rates about $4.20 (about 17.9x less), and Flash at off-peak rates $2.10 (about 36x cheaper).

One agent run: 10M input + 1M output tokens, before cache hits In USD per run

Agent run: $75 on Opus 5, $2.10 on Flash off-peak (One agent run: 10M input + 1M output tokens, before cache hits)
ItemValueClaim
Claude Opus 575 USD per run8
V4.1 Flash, peak rates4.2 USD per run27
V4.1 Flash, off-peak rates2.1 USD per run8

What happened

  • DeepSeek put V4.1 Flash on its official API on September 10, 2026, under the model name deepseek-flash.
  • The model is a 552B-parameter mixture of experts that activates about 8B parameters per token on input and about 16B on output, in a new causal encoder-decoder design.
  • By the post's reading of DeepSeek's own table, Flash ties frontier models on agentic coding and falls well behind on HLE reasoning and on ProgramBench.
  • Serving full precision takes about 614 GB of GPU memory, and day-one recipes from vLLM, SGLang and NVIDIA start at four Blackwell-class GPUs or eight H200s.

Why it matters

  • decision Model choice becomes a per-task routing call: the post sends loops a test suite or type checker can verify to Flash and keeps architecture decisions and ambiguous specs on the frontier model.
  • exposure Callers on unpinned model aliases can be switched to a different model with no code change, as happened when DeepSeek pointed deepseek-v4-pro traffic at Flash indefinitely.
  • exposure Flash agent loops need a sandbox with no reach into real system files; DeepSeek's own report says its agents deleted system files while exploiting vulnerabilities in training.
  • capability The MIT-licensed weights let a team leave DeepSeek's API for OpenRouter or its own servers; the post argues that exit route caps what any host can charge.

I think the cache cut is the best engineering in this release. At 890 bytes per token, a session that fills the 1M-token window holds about 0.89 GB of KV cache [5][26]. At V4 Flash's 3,514 bytes per token, the same token count would take about 3.5 GB, close to four times as much [25][26]. The dev.to post that ran these numbers names the resource it saves: "Long-context agents die on memory, not FLOPs." [12]

The 36x figure comes from a different place: the price list. For 10M input and 1M output tokens, Opus 5 at $5 in and $25 out costs 10 x $5 + 1 x $25 = $75 [7][8]. Flash at its off-peak rate costs 10 x $0.15 + 1 x $0.60 = $2.10, a ratio of 35.7 [8][24]. The post argues the smaller cache is what makes 1M-token sessions "economically sane," but it does not show how DeepSeek's serving costs set its prices [23]. Both bills price every input token at the uncached rate. The post says cache hits are most of an agent loop's input and cost "a rounding error" [18].

Two conditions on usage decide whether 36x reaches an invoice. Peak hours are 01:00-04:00 and 06:00-10:00 UTC on weekdays, seven hours a day, and the post calls moving batch jobs off-peak "Half price for changing a cron time." [9][29] At peak, the same run costs $4.20, about 18x under Opus [27]. Going from low to max effort burns roughly 2.5x the tokens. The post puts the gap at about 15x if everything runs at max [10]. Combine peak hours with max effort, hold the Opus bill at $75, and the Flash run costs $10.50, about 7x less [28].

Then Flash has to do the work. The scores are DeepSeek's own, with no independent verification and no confidence intervals [1]. Across eight agent frameworks, the same checkpoint scored between 65.5 and 74.2 on DeepSWE, and 74.2 is the best case [11]. "Your scaffold can matter more than your model," the post says [19]. Its fix is an afternoon of evals on your own repo, in your own harness, before trusting 74.2 [21].

Compliance is a separate gate. DeepSeek is heading to a Shanghai STAR Market listing, and the post expects teams with views on Chinese-hosted APIs to hold them here too [22]. Self-hosting takes the hosted API out of the picture but leaves the hardware bill. Q4 still needs about 458 GB of VRAM, and a 2-bit quant fits a 128 GB Mac Studio with the quality loss that implies [15].

What to watch

  • Independent DeepSWE runs of V4.1 Flash with confidence intervals, to test DeepSeek's self-reported best-case 74.2.
  • Prices from third-party hosts serving the MIT weights, to test the post's claim that open weights cap hosted prices.
  • Any change to DeepSeek's off-peak discount, its peak windows, or the deepseek-v4-pro redirect to Flash.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence35
Adoption
Insufficient
Hype gap+35
Incentives
Insufficient
Confidence35
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    The benchmark numbers are DeepSeek's own; launch coverage says none are independently verified, and there are no confidence intervals.

  2. [2]

    According to the post's reading of DeepSeek's reported numbers, Flash ties the frontier on agentic coding and gets crushed on hard reasoning (HLE) and on ProgramBench.

    ReportedSupportedSource: dev.to post, reading DeepSeek's reported benchmark table2 sources— create a free account to open themView cited source
  3. [3]

    DeepSeek V4.1 Flash went live on DeepSeek's official API on September 10, 2026, with the API model name deepseek-flash.

    ReportedSupportedSource: dev.to postView cited source

Sources

1 independent publisher whose own reporting we read for this story.

  1. dev.to

    1 article · October 9, 2026

    DeepSeek V4.1 Flash Costs 36x Less Than Claude Opus 5 and Matches It on SWE Benchmarks. Why Is Nobody Panicking?

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Loading related stories