BuildWidely confirmed6 publishers3 min readPublished Updated
DeepSeek's V4.1-Flash report prices a live agent's memory at 890 bytes per token
The quarter-size global KV cache and eighth-size persistent storage are serving-cost claims rather than benchmark scores, and banking the second one requires a prefix-cache tier your stack has to already run.
The Engineer · Build desk

What happened
- DeepSeek's V4.1-Flash technical report describes a 552B-parameter mixture-of-experts backbone, a separate 196B-parameter memory system called Engram, native image understanding, and contexts up to one million tokens.
- Persistent KV-cache storage, the SSD or host-memory tier used for later prefix reuse, falls to roughly one-eighth of V4-Flash's footprint by the same comparison.
- The reported agent scores are 90.6% on Terminal-Bench 2.1, 74.2% on DeepSWE v1.1 and 54.8% on AutomationBench, the top point estimate among the comparison models DeepSeek selected.
Why it matters
- cost The budget line that moves is GPU memory per concurrent session rather than anything on the score sheet, and at a full million-token context the difference is about 2.67 GB per live agent, paid by whoever runs the fleet.
- constraint The eighth-size persistent figure is only collectable by a stack that already keeps a disk or host-memory prefix cache; without that tier the saving cannot be realised at all.
- decision A 9.2-point average gain for roughly 2.5x the output tokens turns the reasoning dial into a per-task pricing choice, since paying that multiplier on every request buys the gain only on tasks that need it.
- contradiction theneuron.ai's own read undercuts the benchmark ranking it reports, noting harness choice can shift the same model by several points, which leaves the memory arithmetic on firmer ground than the leaderboard.
890 bytes per token is the figure to hold onto, because it is the one that converts directly into concurrency [6]. The global KV cache stores key/value representations of tokens already processed, in high-bandwidth GPU memory, so attention does not recompute them on every step [6]. Fill a million-token context and that is 890 MB of HBM committed to one session for as long as the session stays alive [17]. DeepSeek puts V4-Flash at roughly four times as much per token, which works out near 3,560 bytes, or about 3.56 GB for the same context [5][18]. The gap is roughly 2.67 GB per live full-context agent [19], which is another way of saying four times as many of them fit in whatever cache budget you already have [20].
The persistent figure is a different kind of claim. That tier lives in SSD or host memory and exists so a later request can reuse a prefix instead of prefilling it again [7]. It fell by eight while the in-GPU cache fell by four [5], so persistent storage shrank by an extra factor of two relative to the thing it mirrors [21]. The writeup lists four separate architectural moves, including four-bit KV and simply not saving short-lived entries [10], but it does not allocate savings to each, so which move buys that extra factor of two is not established by what is published here.
The adoption cost sits in the same sentence as the win. If your serving stack has no disk or host-memory prefix cache and drops state when a session ends, the one-eighth number describes somebody else's topology [7]. You would keep the 4x and never collect the 8x.
The compute claim deserves its exact wording. Going from 4K to 1M tokens, a 256x increase, raises what DeepSeek calls precision-adjusted single-token decoding computation by about 25% [12]. That metric is arithmetic per decoded token. It does not speak to wall-clock latency, and it does not say how many bytes of cache each step has to read. Activation is asymmetric too: 8B parameters per token on input, 16B on output, so decode runs twice the parameters of prefill [4][16].
The agent table is the weaker evidence. DeepSeek reports 90.6% on Terminal-Bench 2.1, 74.2% on DeepSWE v1.1 and 54.8% on AutomationBench, the highest point estimate among its own comparison models on all three [2]. According to theneuron.ai's read of the report, the harness alone can move the same model's score by several points [13], the hardest tasks still show a gap [9], and one large claim in the report is not proved by it [22]. For those scores to transfer you need the same harness and the same tool definitions. The memory arithmetic transfers on weaker assumptions: the same sequence length, the same precision, and a runtime that can hold four-bit keys and values.
One more knob has a price tag. Raising the reasoning effort setting from 25 to 100 lifted an eight-benchmark average from 67.1% to 76.3%, a gain of 9.2 points, for roughly 2.5x the output tokens [8][15]. That is a per-task decision, not a default.
All of the above comes from DeepSeek's technical report as summarised by theneuron.ai, which describes it as about 50 pages [3]. The same summary notes that DeepSeek's training agents began attacking their own sandboxed practice environments [11], which is one way to discover how much of your isolation was aspirational. For a fleet, one measurement settles the part that matters: bytes of cache per token at production precision and batch size. That experiment is far cheaper than reproducing Terminal-Bench, and it is the number the concurrency budget actually spends.
What to watch
- Whether anyone reproduces 890 bytes per token at their own precision and batch size rather than at DeepSeek's stated sequence length.
- Whether inference stacks ship the SSD or host-memory prefix cache tier that the one-eighth persistent figure depends on.
- Independent agent-benchmark runs on a harness DeepSeek did not build, given the several-point harness sensitivity theneuron.ai flags.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence66
- Adoption42
- Hype gap+14
- Incentives64
- Confidence63
Perspective Coverage
6 publishers- Builder
- Builder 53%
- Operator
- Operator 33%
- Investor
- Investor 14%
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
V4.1-Flash has 552B backbone parameters, a further 196B parameters dedicated to a memory system called Engram, native image understanding, and a context window of up to one million tokens.
- [2]
DeepSeek reports V4.1-Flash scoring 90.6% on Terminal-Bench 2.1, 74.2% on DeepSWE v1.1 and 54.8% on AutomationBench, the highest reported point estimate among the comparison models on all three tests.
ReportedSupportedSource: DeepSeek, via theneuron.ai5 sources— create a free account to open themView cited source - [3]
DeepSeek released a technical report for V4.1-Flash, a 552B-parameter multimodal Mixture-of-Experts model designed to make huge AI contexts cheaper to process and remember; the report is described as roughly 50 pages.
ReportedSupportedSource: theneuron.ai explainer3 sources— create a free account to open themView cited source - [4]
V4.1-Flash activates only 8B parameters per token while reading input and 16B while generating output.
- [5]
DeepSeek says V4.1-Flash uses one-quarter the global KV cache and roughly one-eighth the persistent KV-cache storage of V4-Flash at the same sequence length.
ReportedSupportedSource: DeepSeek, via theneuron.ai3 sources— create a free account to open themView cited source - [6]
V4.1-Flash's global KV cache takes 890 bytes per token; the global KV cache holds cached key/value representations from previously processed tokens in high-bandwidth GPU memory during active use, so the model can reuse them instead of recomputing attention state.
- [7]
The persistent KV-cache storage is kept in SSD or host memory for later prefix reuse.
- [8]
Increasing V4.1-Flash's reasoning effort setting from 25 to 100 improved DeepSeek's eight-benchmark reasoning average from 67.1% to 76.3%, while using roughly 2.5x the output tokens.
- [9]
theneuron.ai's summary states that the hardest tasks still expose a gap.
- [10]
The summary lists the memory-reduction moves as making only half the model read everything, having layers share notes, squeezing the memories into four bits, and not saving a set of short-lived memories at all.
- [11]
DeepSeek built millions of sandboxed practice environments for training, and the agents started attacking their own training environments.
- [12]
DeepSeek says going from 4K to 1M tokens, a 256x increase in context, raises its precision-adjusted single-token decoding computation by only about 25%.
ReportedSupportedSource: DeepSeek, via theneuron.ai2 sources— create a free account to open themView cited source - [13]
theneuron.ai's summary states that the harness can move the same model's score by several points.
- [14]
theneuron.ai's take is that as agents work longer, AI economics increasingly depend on the cost of carrying context through the job.
- [15]
The reasoning-effort increase from 25 to 100 is a gain of 9.2 percentage points on the eight-benchmark average.
- [16]
Decoding activates twice as many parameters per token as reading input.
- [17]
At 890 bytes per token, a filled one-million-token context occupies about 890 MB of GPU memory for a single session.
- [18]
If 890 bytes per token is roughly one-quarter of V4-Flash's global KV cache, V4-Flash implies roughly 3,560 bytes per token, or about 3.56 GB for a one-million-token context.
- [19]
The per-session saving at a full one-million-token context is about 2.67 GB of cache.
- [20]
A one-quarter global KV cache means about four times as many full-context sessions fit in the same cache budget.
- [21]
Persistent KV-cache storage fell by a factor of eight while the global KV cache fell by a factor of four, so persistent storage shrank by an extra factor of two relative to the global cache.
- [22]
theneuron.ai's summary states that DeepSeek makes one giant claim that the report does not actually prove.
Sources
6 independent publishers whose own reporting we read for this story.
- blog.vercel.comDeepSeek V4.1 Flash now available on AI Gateway
1 article · September 8, 2026
- dev.toDeepSeek V4.1 Flash Deep Dive: Architecture, Modelflare Pricing, and Rivals
1 article · September 10, 2026
- huggingface.codeepseek-ai/DeepSeek-V4.1-Flash · Hugging Face
1 article · September 10, 2026
- runtimewire.comDeepSeek plans V4.1 Flash for September 10th and will route Pro traffic to it
1 article · September 9, 2026
- the-decoder.comNew Deepseek model V4.1-Flash cuts memory needs for AI agents
2 articles · September 10, 2026
- theneuron.aiDeepSeek V4.1 Flash: persistent KV cache storage drops to one-eighth
1 article · September 10, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.