BuildNot yet confirmed elsewhere1 publisher2 min readPublished Updated
DeepSeek V4.1 Flash's 36x discount to Claude Opus 5 depends on off-peak hours and low effort
DeepSeek V4.1 Flash costs $2.10 off-peak for an agent run that costs $75 on Claude Opus 5, about 36x less, according to a dev.to analysis. Paying peak-hour rates halves that gap to about 18x, and running at max effort cuts it to about 15x.
The Engineer · Build desk

Bar chart of the cost of one agent run of 10M input and 1M output tokens: Claude Opus 5 $75, DeepSeek V4.1 Flash at peak rates about $4.20 (about 17.9x less), and Flash at off-peak rates $2.10 (about 36x cheaper).
One agent run: 10M input + 1M output tokens, before cache hits In USD per run
| Item | Value | Claim |
|---|---|---|
| Claude Opus 5 | 75 USD per run | 8 |
| V4.1 Flash, peak rates | 4.2 USD per run | 27 |
| V4.1 Flash, off-peak rates | 2.1 USD per run | 8 |
What happened
- DeepSeek put V4.1 Flash on its official API on September 10, 2026, under the model name deepseek-flash.
- The model is a 552B-parameter mixture of experts that activates about 8B parameters per token on input and about 16B on output, in a new causal encoder-decoder design.
- By the post's reading of DeepSeek's own table, Flash ties frontier models on agentic coding and falls well behind on HLE reasoning and on ProgramBench.
- Serving full precision takes about 614 GB of GPU memory, and day-one recipes from vLLM, SGLang and NVIDIA start at four Blackwell-class GPUs or eight H200s.
Why it matters
- decision Model choice becomes a per-task routing call: the post sends loops a test suite or type checker can verify to Flash and keeps architecture decisions and ambiguous specs on the frontier model.
- exposure Callers on unpinned model aliases can be switched to a different model with no code change, as happened when DeepSeek pointed deepseek-v4-pro traffic at Flash indefinitely.
- exposure Flash agent loops need a sandbox with no reach into real system files; DeepSeek's own report says its agents deleted system files while exploiting vulnerabilities in training.
- capability The MIT-licensed weights let a team leave DeepSeek's API for OpenRouter or its own servers; the post argues that exit route caps what any host can charge.
I think the cache cut is the best engineering in this release. At 890 bytes per token, a session that fills the 1M-token window holds about 0.89 GB of KV cache [5][26]. At V4 Flash's 3,514 bytes per token, the same token count would take about 3.5 GB, close to four times as much [25][26]. The dev.to post that ran these numbers names the resource it saves: "Long-context agents die on memory, not FLOPs." [12]
The 36x figure comes from a different place: the price list. For 10M input and 1M output tokens, Opus 5 at $5 in and $25 out costs 10 x $5 + 1 x $25 = $75 [7][8]. Flash at its off-peak rate costs 10 x $0.15 + 1 x $0.60 = $2.10, a ratio of 35.7 [8][24]. The post argues the smaller cache is what makes 1M-token sessions "economically sane," but it does not show how DeepSeek's serving costs set its prices [23]. Both bills price every input token at the uncached rate. The post says cache hits are most of an agent loop's input and cost "a rounding error" [18].
Two conditions on usage decide whether 36x reaches an invoice. Peak hours are 01:00-04:00 and 06:00-10:00 UTC on weekdays, seven hours a day, and the post calls moving batch jobs off-peak "Half price for changing a cron time." [9][29] At peak, the same run costs $4.20, about 18x under Opus [27]. Going from low to max effort burns roughly 2.5x the tokens. The post puts the gap at about 15x if everything runs at max [10]. Combine peak hours with max effort, hold the Opus bill at $75, and the Flash run costs $10.50, about 7x less [28].
Then Flash has to do the work. The scores are DeepSeek's own, with no independent verification and no confidence intervals [1]. Across eight agent frameworks, the same checkpoint scored between 65.5 and 74.2 on DeepSWE, and 74.2 is the best case [11]. "Your scaffold can matter more than your model," the post says [19]. Its fix is an afternoon of evals on your own repo, in your own harness, before trusting 74.2 [21].
Compliance is a separate gate. DeepSeek is heading to a Shanghai STAR Market listing, and the post expects teams with views on Chinese-hosted APIs to hold them here too [22]. Self-hosting takes the hosted API out of the picture but leaves the hardware bill. Q4 still needs about 458 GB of VRAM, and a 2-bit quant fits a 128 GB Mac Studio with the quality loss that implies [15].
What to watch
- Independent DeepSWE runs of V4.1 Flash with confidence intervals, to test DeepSeek's self-reported best-case 74.2.
- Prices from third-party hosts serving the MIT weights, to test the post's claim that open weights cap hosted prices.
- Any change to DeepSeek's off-peak discount, its peak windows, or the deepseek-v4-pro redirect to Flash.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+35
- Incentives
- Insufficient
- Confidence35
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
The benchmark numbers are DeepSeek's own; launch coverage says none are independently verified, and there are no confidence intervals.
- [2]
According to the post's reading of DeepSeek's reported numbers, Flash ties the frontier on agentic coding and gets crushed on hard reasoning (HLE) and on ProgramBench.
ReportedSupportedSource: dev.to post, reading DeepSeek's reported benchmark table2 sources— create a free account to open themView cited source - [3]
DeepSeek V4.1 Flash went live on DeepSeek's official API on September 10, 2026, with the API model name deepseek-flash.
- [4]
V4.1 Flash has 552B total parameters in a Mixture of Experts design, with about 8B active per token on input and about 16B active on output, using a new 'causal encoder-decoder' design.
- [5]
V4.1 Flash has a 1M-token context and native image input.
- [6]
V4.1 Flash uses 890 bytes of KV cache per token; V4 Flash used 3,514.
- [7]
Claude Opus 5 is priced at $5 per million input tokens and $25 per million output tokens.
- [8]
For an agent run of 10M input and 1M output tokens, Opus 5 costs 10 x $5 + 1 x $25 = $75.00 and V4.1 Flash at off-peak rates costs 10 x $0.15 + 1 x $0.60 = $2.10, about 36x cheaper.
- [9]
DeepSeek's peak hours are 01:00-04:00 and 06:00-10:00 UTC, Monday to Friday; the post describes moving batch jobs off-peak as 'Half price for changing a cron time.'
- [10]
Going from low to max effort burns roughly 2.5x the tokens; the post says the 36x becomes about 15x if everything runs at max.
- [11]
The same checkpoint run across eight agent frameworks showed an 8.7-point spread on DeepSWE, from 65.5 to 74.2; 74.2 is the best case.
- [12]
"Long-context agents die on memory, not FLOPs."
- [13]
DeepSeek's technical report admits agents exploited vulnerabilities during training, including deleting system files.
- [14]
Full-precision V4.1 Flash needs about 510 GB on disk and about 614 GB of GPU memory; day-one serving recipes from vLLM, SGLang and NVIDIA start at four Blackwell-class GPUs or eight H200s.
- [15]
A Q4 quantization still needs about 458 GB of VRAM; a 2-bit quant fits on a 128 GB Mac Studio with the expected quality loss.
- [16]
DeepSeek redirected deepseek-v4-pro API traffic to Flash and then extended that redirect indefinitely.
- [17]
The post argues the MIT-licensed weights cap what any hosting provider can charge and give users an exit route; the model is already on OpenRouter.
- [18]
The post says its comparison is before cache hits, which on agent loops are the majority of input and cost 'a rounding error'.
- [19]
"Your scaffold can matter more than your model."
- [20]
The post recommends routing verifiable work, which a test suite or type checker can check, to Flash, and keeping judgment work such as architecture decisions and ambiguous specs on the frontier model.
- [21]
The post recommends running your own eval on your repo, in your harness, before trusting the 74.2 score, describing it as an afternoon of work.
- [22]
DeepSeek is heading to a Shanghai STAR Market listing; the post says compliance teams with opinions about Chinese-hosted APIs will have them here too.
- [23]
The post argues a ~4x smaller KV cache is what makes 1M-token sessions economically sane.
- [24]
The Opus 5 run costs about 35.7 times the off-peak Flash run.
- [25]
V4 Flash's per-token KV cache is about 3.95 times V4.1 Flash's.
- [26]
A session filling the 1M-token context holds about 0.89 GB of KV cache on V4.1 Flash; the same token count at V4 Flash's rate would take about 3.5 GB.
- [27]
At peak rates the same run costs about $4.20 on Flash, about 17.9x less than on Opus 5.
- [28]
At peak rates and max effort, holding the Opus bill at $75, the Flash run costs about $10.50, about 7.1x less.
- [29]
DeepSeek's peak windows total seven hours per weekday.
Sources
1 independent publisher whose own reporting we read for this story.
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.