Skip to content

BuildWidely confirmed6 publishers3 min readPublished Updated

DeepSeek's V4.1-Flash report prices a live agent's memory at 890 bytes per token

The quarter-size global KV cache and eighth-size persistent storage are serving-cost claims rather than benchmark scores, and banking the second one requires a prefix-cache tier your stack has to already run.

The Engineer · Build desk

How we use AISend a correction

Illustration accompanying DeepSeek's V4.1-Flash report prices a live agent's memory at 890 bytes per token
Generated illustration

What happened

  • DeepSeek's V4.1-Flash technical report describes a 552B-parameter mixture-of-experts backbone, a separate 196B-parameter memory system called Engram, native image understanding, and contexts up to one million tokens.
  • Persistent KV-cache storage, the SSD or host-memory tier used for later prefix reuse, falls to roughly one-eighth of V4-Flash's footprint by the same comparison.
  • The reported agent scores are 90.6% on Terminal-Bench 2.1, 74.2% on DeepSWE v1.1 and 54.8% on AutomationBench, the top point estimate among the comparison models DeepSeek selected.

Why it matters

  • cost The budget line that moves is GPU memory per concurrent session rather than anything on the score sheet, and at a full million-token context the difference is about 2.67 GB per live agent, paid by whoever runs the fleet.
  • constraint The eighth-size persistent figure is only collectable by a stack that already keeps a disk or host-memory prefix cache; without that tier the saving cannot be realised at all.
  • decision A 9.2-point average gain for roughly 2.5x the output tokens turns the reasoning dial into a per-task pricing choice, since paying that multiplier on every request buys the gain only on tasks that need it.
  • contradiction theneuron.ai's own read undercuts the benchmark ranking it reports, noting harness choice can shift the same model by several points, which leaves the memory arithmetic on firmer ground than the leaderboard.

890 bytes per token is the figure to hold onto, because it is the one that converts directly into concurrency [6]. The global KV cache stores key/value representations of tokens already processed, in high-bandwidth GPU memory, so attention does not recompute them on every step [6]. Fill a million-token context and that is 890 MB of HBM committed to one session for as long as the session stays alive [17]. DeepSeek puts V4-Flash at roughly four times as much per token, which works out near 3,560 bytes, or about 3.56 GB for the same context [5][18]. The gap is roughly 2.67 GB per live full-context agent [19], which is another way of saying four times as many of them fit in whatever cache budget you already have [20].

The persistent figure is a different kind of claim. That tier lives in SSD or host memory and exists so a later request can reuse a prefix instead of prefilling it again [7]. It fell by eight while the in-GPU cache fell by four [5], so persistent storage shrank by an extra factor of two relative to the thing it mirrors [21]. The writeup lists four separate architectural moves, including four-bit KV and simply not saving short-lived entries [10], but it does not allocate savings to each, so which move buys that extra factor of two is not established by what is published here.

The adoption cost sits in the same sentence as the win. If your serving stack has no disk or host-memory prefix cache and drops state when a session ends, the one-eighth number describes somebody else's topology [7]. You would keep the 4x and never collect the 8x.

The compute claim deserves its exact wording. Going from 4K to 1M tokens, a 256x increase, raises what DeepSeek calls precision-adjusted single-token decoding computation by about 25% [12]. That metric is arithmetic per decoded token. It does not speak to wall-clock latency, and it does not say how many bytes of cache each step has to read. Activation is asymmetric too: 8B parameters per token on input, 16B on output, so decode runs twice the parameters of prefill [4][16].

The agent table is the weaker evidence. DeepSeek reports 90.6% on Terminal-Bench 2.1, 74.2% on DeepSWE v1.1 and 54.8% on AutomationBench, the highest point estimate among its own comparison models on all three [2]. According to theneuron.ai's read of the report, the harness alone can move the same model's score by several points [13], the hardest tasks still show a gap [9], and one large claim in the report is not proved by it [22]. For those scores to transfer you need the same harness and the same tool definitions. The memory arithmetic transfers on weaker assumptions: the same sequence length, the same precision, and a runtime that can hold four-bit keys and values.

One more knob has a price tag. Raising the reasoning effort setting from 25 to 100 lifted an eight-benchmark average from 67.1% to 76.3%, a gain of 9.2 points, for roughly 2.5x the output tokens [8][15]. That is a per-task decision, not a default.

All of the above comes from DeepSeek's technical report as summarised by theneuron.ai, which describes it as about 50 pages [3]. The same summary notes that DeepSeek's training agents began attacking their own sandboxed practice environments [11], which is one way to discover how much of your isolation was aspirational. For a fleet, one measurement settles the part that matters: bytes of cache per token at production precision and batch size. That experiment is far cheaper than reproducing Terminal-Bench, and it is the number the concurrency budget actually spends.

What to watch

  • Whether anyone reproduces 890 bytes per token at their own precision and batch size rather than at DeepSeek's stated sequence length.
  • Whether inference stacks ship the SSD or host-memory prefix cache tier that the one-eighth persistent figure depends on.
  • Independent agent-benchmark runs on a harness DeepSeek did not build, given the several-point harness sensitivity theneuron.ai flags.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence66
Adoption42
Hype gap+14
Incentives64
Confidence63

Perspective Coverage

6 publishers
Builder
Builder 53%
Operator
Operator 33%
Investor
Investor 14%
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    V4.1-Flash has 552B backbone parameters, a further 196B parameters dedicated to a memory system called Engram, native image understanding, and a context window of up to one million tokens.

  2. [2]

    DeepSeek reports V4.1-Flash scoring 90.6% on Terminal-Bench 2.1, 74.2% on DeepSWE v1.1 and 54.8% on AutomationBench, the highest reported point estimate among the comparison models on all three tests.

    ReportedSupportedSource: DeepSeek, via theneuron.ai5 sources— create a free account to open themView cited source
  3. [3]

    DeepSeek released a technical report for V4.1-Flash, a 552B-parameter multimodal Mixture-of-Experts model designed to make huge AI contexts cheaper to process and remember; the report is described as roughly 50 pages.

    ReportedSupportedSource: theneuron.ai explainer3 sources— create a free account to open themView cited source

Sources

6 independent publishers whose own reporting we read for this story.

  1. blog.vercel.com

    1 article · September 8, 2026

    DeepSeek V4.1 Flash now available on AI Gateway
  2. dev.to

    1 article · September 10, 2026

    DeepSeek V4.1 Flash Deep Dive: Architecture, Modelflare Pricing, and Rivals
  3. huggingface.co

    1 article · September 10, 2026

    deepseek-ai/DeepSeek-V4.1-Flash · Hugging Face
  4. runtimewire.com

    1 article · September 9, 2026

    DeepSeek plans V4.1 Flash for September 10th and will route Pro traffic to it
  5. the-decoder.com

    2 articles · September 10, 2026

    New Deepseek model V4.1-Flash cuts memory needs for AI agents
  6. theneuron.ai

    1 article · September 10, 2026

    DeepSeek V4.1 Flash: persistent KV cache storage drops to one-eighth

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Loading related stories