Skip to content

Build1 publisher3 min readPublished

Agent memory handoffs cut replaced context 98.65% by moving verbatim text out of the prompt

Memory handoffs cut replaced agent context by 98.65% across 10,241 completed runs, according to an audit published on dev.to. For agent builders, the open risk now sits in whether the agent recalls the right stored text, since lossy summaries are no longer the failure point.

The Engineer · Build desk

What happened

  • The author recomputed hashes for all 416,472 retained segments and found none missing exact source text and no hash mismatches.
  • Segment counters recorded 416,459 KEEP decisions, 13 COMPRESS and no DROP, and compression shortened stored text by only 818 characters.
  • Each handoff holds a short navigation summary, a manifest identifier, selected content hashes and instructions for fetching the omitted text.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint The 98.65% cannot be dropped into a token or cost forecast, because system instructions, tool definitions, protected recent messages and recall results still fill context.
  • decision Adopting the design commits a team to building and evaluating recall, because in the test the correct answers depended on what retrieval fetched into context.
  • exposure Keeping nearly every segment verbatim puts the full conversation history in an external store, and redacting secrets before persistence is the one control on it the post describes.

The per-run averages make the headline figure concrete. Each replaced window held about 111,062 characters, and the handoff that replaced it about 1,501 [1]. According to the author of the audit, the 98.65% comes from summed lengths across all runs, not from an average of per-run percentages, so the longest windows count most [1][2].

Here is what happens to a window before any of that saving shows up, per the post:

1. The agent selects a message-safe middle window and cuts its messages into role-aware segments of at most 3,800 characters [8]. 2. Tool, system and developer segments are protected, along with recognized patterns such as code fences, commands, paths, hashes and error evidence. These resolve to KEEP before any model classifies them [9]. 3. Everything else goes to a GLiClass router. Weak decisions and model failures fall back to KEEP [10]. 4. COMPRESS candidates pass through LLMLingua and then an NLI gate [10].

Every default in that sequence leans toward keeping text. Compression ended up handling about 0.003% of segments [2]. A stage that fires that rarely is, for accounting purposes, a KEEP stage with extra dependencies. The context saving comes from the handoff itself: once a window's retained material is durably stored, the whole window leaves the active context, KEEP segments included [6]. "The large active-context change comes from externalization," the author wrote [7].

I think falling back to KEEP is the right default for agent work. A misrouted classifier costs storage. A misrouted compressor can cost the exact port a later tool call needs. The planner smoke test in the deployed API environment exercised that case: 24 synthetic tool messages, each with an exact endpoint, port and command, were all retained verbatim and still yielded a much smaller handoff [17]. That test made no database writes, and its timing was a single observation [17].

The hash audit establishes that the stored text is intact. Whether an agent can use it has been tested once, on synthetic data with a live answering model and real embeddings [11]. Set aside the question whose correct answer was UNKNOWN, and the handoff-only arm scored zero of 11 [3]. The recall arm matched full history while using substantially fewer input tokens, retrieved evidence included [12]. "This is a small controlled result, not a claim of general lossless agent memory," the author wrote [14].

Repeated runs can contain related material, so the 1.14 billion source characters are not a count of unique information [1][16]. For the 98.65% to carry over to another deployment, its windows would need a similar size profile, and its agent would need to decide for itself when to recall. In the synthetic test that decision belonged to the harness, which made the retrieval calls [13].

What to watch

  • A rerun of the recall comparison in which the agent, not the harness, decides when and what to retrieve, with its answer rate.
  • A whole-prompt token or inference-cost measurement that counts recall results, which would say whether the character saving survives as a billing saving.
  • A latency distribution for the planner in the deployed API, since the smoke test recorded a single timing.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories