Build1 distinct publisher3 min readUpdated
IBM Research ran self-mined guidelines across eight models on AppWorld. One model gained 16.1 points for 5 percent more tokens; another gained nothing at all.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
IBM Research, writing on the Hugging Face blog, has published an eight-model sweep of ALTK-Evolve, a loop that mines behavioural guidelines from an agent's own past runs and injects them back at inference time [1][2]. The finding that matters for anyone budgeting context: the same guideline set that moved one model 16.1 percentage points moved another zero, and the authors frame memory as a dose to calibrate rather than a feature to enable [3][4].
"Memory" here is not transcript replay. It is a distilled set of strategies that worked, mistakes to avoid, and edge cases, extracted from both successful and unsuccessful trajectories, consolidated, and then served either whole or as a task-relevant subset [5][6]. No weights are updated and no human annotation is involved, which is what makes it portable across models [7]. Evaluation ran on AppWorld: 585 multi-step tasks split 168 test_normal and 417 test_challenge across nine simulated apps, scored on task goal completion and the stricter all-or-nothing scenario goal completion [8][9][10].
Three patterns emerged. DeepSeek-V3.2, a 671B mixture-of-experts model, gained 9.5 points of task completion from its full self-mined guideline set [11]. gpt-oss-120b, a 117B MoE, did best on a compact high-confidence core plus a handful of guidelines retrieved per task, gaining 16.1 points at roughly 5 percent more tokens; its full guideline set gained less and cost about 50 percent more tokens [12][13]. GLM-5, at 745B, showed no measurable gain, a state the authors label "saturated" while explicitly saying the label describes the observation and not a cause [14][15].
The arithmetic is the interesting part. The curated configuration returned about 3.2 points of task completion per additional percent of tokens on gpt-oss-120b [16]. If both token figures are read against the same no-memory baseline, the full set costs roughly ten times the token premium of curated retrieval on that model, for a smaller gain [17]. And the largest gain in the three named results came from the smallest of the three models [18], which is consistent with the authors' warning that parameter count does not predict which pattern a model lands in; they point instead to benchmark headroom, context-window size, architecture, guideline quality and task distribution, and say separating those factors is still open work [19].
Two methodological details deserve credit. The guideline set was mined once from AppWorld's training split only, with no test-split data entering its construction, and the two configurations differ only in delivery, not in provenance [20]. And because the number of guidelines a model mines depends on its own capability, results are reported by strategy rather than by raw guideline counts, which the authors say are not comparable across models [21]. That also means you cannot read a token budget off this work directly; the authors note prompt caching keeps even the full set affordable in production, which shifts where the cost actually lands [22].
What to watch: whether anyone publishes the scenario-level numbers alongside task completion, since the stricter bar is where guideline noise should show up first [10]; whether the saturated pattern survives on a task distribution with more headroom than AppWorld [15]; and whether a cheap pre-flight test emerges that tells you which of the three doses a given model wants before you pay for the full sweep [4].
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
The authors state that agentic memory is not a feature you switch on but a dose you calibrate to the model, and that the right dose depends on the model and can be calibrated.
ALTK-Evolve lets an agent learn from its own past trajectories by distilling reusable guidelines and injecting them back at inference time, with no weight updates and no human annotation.
The evaluation was scaled to eight models, from a 30B dense model to frontier proprietary systems.
Across the eight models, gains ranged from a 16.1 percentage point improvement in task completion (gpt-oss-120b, curated retrieval) to no measurable gain (GLM-5).
"Memory" in this work does not mean replaying a past transcript; it means a guideline set of strategies that worked, mistakes to avoid, and edge cases, distilled from the agent's own prior trajectories.
The authors say what puts a model into one pattern is not simply parameter count; benchmark headroom, context-window size, architecture, guideline quality and task distribution all appear to shape it, and separating those factors is ongoing work.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Detailed but self-reported and single-source
The claims are unusually specific for a vendor post: a named public benchmark with task counts and splits, two scoring metrics, an explicit train-split-only mining protocol, per-model deltas on both TGC and SGC, and token overhead measured against a no-memory baseline. Weighing against that, everything comes from one IBM Research blog post with no independent replication, no visible code or paper artifact, results for only five of eight swept models in the readable text, truncated results and token tables, and the authors' own admission that the causal drivers of the observed patterns are unseparated.
No adoption signal in supplied sources
The only observations available are the authors' own benchmark runs and token measurements. The supplied material contains no third-party deployment, download, integration, customer, or production-usage disclosure for ALTK-Evolve, and the authors' portability claim covers models tested rather than users adopting. Inferring adoption from a self-run benchmark would be a guess.
Mildly overstated framing over hedged data
The post's generalizing frame — a dose-response curve for agentic memory, tiers of models, 'the right dose depends on the model and we can calibrate it' — runs ahead of what one lab's runs on one simulated-app benchmark can establish, and the production-affordability claim is asserted rather than measured. The gap stays small because the authors hedge conspicuously: they publish a zero-gain model, state that 'saturated' names an observation and not a cause, decline to attribute placement to parameter count, and disclose that the full guideline set underperformed the cheaper option on gpt-oss-120b.
Originating lab publishing its own toolkit results
Every figure in the cluster comes from IBM Research evaluating IBM Research's own ALTK-Evolve, published on a model-hub blog that serves as a distribution and recruitment channel for the toolkit. Framing choices favor the method: a headline efficiency win, an x-axis starting at 40 percent to make bars visible, and the mining pass's own cost left out of the reported overhead. The score is not higher because the same post volunteers a null result, flags the limits of its own labels, and specifies a contamination control.
Moderate-low: one publisher, one truncated source
Confidence is limited by structure rather than by internal contradiction. There is a single publisher and a single source item, the body is truncated mid-sentence so the results and token tables cannot be read, three of eight models are named as pattern exemplars, and no independent evaluation exists in the cluster to corroborate or contest the figures. The specificity and internal consistency of what is shown — including the split arithmetic checking out — keeps confidence from falling further.
leadership
Cost per successful task, not per token: a 2,400-run benchmark reorders the model shortlist1 distinct publisher
science
GLM-5.3 says the quiet part: the base model did not change, the post-training did1 distinct publisher
build
GLM-5.3 is a paper, not an endpoint: Z.ai publishes research before weights1 distinct publisher
build
Inco AI's DFlash 2: 21% longer accepted drafts for 1.3% latency and 18.5M parameters1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 18, 2026