Skip to content

Build1 publisher3 min readPublished

Agent memory has a dose-response curve, and the cheapest dose won the biggest gain

IBM Research ran self-mined guidelines across eight models on AppWorld. One model gained 16.1 points for 5 percent more tokens; another gained nothing at all.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Agent memory has a dose-response curve, and the cheapest dose won the biggest gain
Generated illustration

What happened

  • ALTK-Evolve lets an agent learn from its own past trajectories by distilling reusable guidelines and injecting them back at inference time, with no weight updates and no human annotation.
  • The evaluation was scaled to eight models, from a 30B dense model to frontier proprietary systems.
  • Across the eight models, gains ranged from a 16.1 percentage point improvement in task completion (gpt-oss-120b, curated retrieval) to no measurable gain (GLM-5).
  • The authors state that agentic memory is not a feature you switch on but a dose you calibrate to the model, and that the right dose depends on the model and can be calibrated.
  • "Memory" in this work does not mean replaying a past transcript; it means a guideline set of strategies that worked, mistakes to avoid, and edge cases, distilled from the agent's own prior trajectories.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

IBM Research, writing on the Hugging Face blog, has published an eight-model sweep of ALTK-Evolve, a loop that mines behavioural guidelines from an agent's own past runs and injects them back at inference time [1][2]. The finding that matters for anyone budgeting context: the same guideline set that moved one model 16.1 percentage points moved another zero, and the authors frame memory as a dose to calibrate rather than a feature to enable [3][4].

"Memory" here is not transcript replay. It is a distilled set of strategies that worked, mistakes to avoid, and edge cases, extracted from both successful and unsuccessful trajectories, consolidated, and then served either whole or as a task-relevant subset [5][6]. No weights are updated and no human annotation is involved, which is what makes it portable across models [7]. Evaluation ran on AppWorld: 585 multi-step tasks split 168 test_normal and 417 test_challenge across nine simulated apps, scored on task goal completion and the stricter all-or-nothing scenario goal completion [8][9][10].

Three patterns emerged. DeepSeek-V3.2, a 671B mixture-of-experts model, gained 9.5 points of task completion from its full self-mined guideline set [11]. gpt-oss-120b, a 117B MoE, did best on a compact high-confidence core plus a handful of guidelines retrieved per task, gaining 16.1 points at roughly 5 percent more tokens; its full guideline set gained less and cost about 50 percent more tokens [12][13]. GLM-5, at 745B, showed no measurable gain, a state the authors label "saturated" while explicitly saying the label describes the observation and not a cause [14][15].

The arithmetic is the interesting part. The curated configuration returned about 3.2 points of task completion per additional percent of tokens on gpt-oss-120b [16]. If both token figures are read against the same no-memory baseline, the full set costs roughly ten times the token premium of curated retrieval on that model, for a smaller gain [17]. And the largest gain in the three named results came from the smallest of the three models [18], which is consistent with the authors' warning that parameter count does not predict which pattern a model lands in; they point instead to benchmark headroom, context-window size, architecture, guideline quality and task distribution, and say separating those factors is still open work [19].

Two methodological details deserve credit. The guideline set was mined once from AppWorld's training split only, with no test-split data entering its construction, and the two configurations differ only in delivery, not in provenance [20]. And because the number of guidelines a model mines depends on its own capability, results are reported by strategy rather than by raw guideline counts, which the authors say are not comparable across models [21]. That also means you cannot read a token budget off this work directly; the authors note prompt caching keeps even the full set affordable in production, which shifts where the cost actually lands [22].

What to watch: whether anyone publishes the scenario-level numbers alongside task completion, since the stricter bar is where guideline noise should show up first [10]; whether the saturated pattern survives on a task distribution with more headroom than AppWorld [15]; and whether a cheap pre-flight test emerges that tells you which of the three doses a given model wants before you pay for the full sweep [4].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories