Build1 distinct publisher3 min readUpdated
Deterministic checkers replaced the judge model, and the project instruction file finished below having no instruction file at all. The measurable part of context engineering turns out to be narrow.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
The 22.2 points is the gap between two pass rates, and it is also the entire paired result: 17 comparisons went to RE-call, one went to the instruction file [6], and 16 net wins across 72 pairs is 22.2 percent [1]. Which leaves 54 of the 72 comparisons where both arms did the same thing, passing together or failing together [2]. The paired McNemar test at p = 0.000145 [3] says the direction is real. It does not say the effect is broad. Three quarters of a suite built around line ending handling and money rounding [14] could not tell the two configurations apart.
Where the effect does live is stated plainly. The eight tasks the author, writing on dev.to, labels memory-sensitive improved by 45.8 points [8], which over 24 cells is 11 net wins [3]. The other sixteen tasks carry the remaining 5 net wins across 48 cells, about 10.4 points [4]. One third of the suite [5] produces roughly two thirds of the headline, and the memory-sensitive label was applied by the person who built the memory layer. Four more tasks in that bucket and the number grows without a line of code changing.
The finding that costs nothing to act on is the one the author says he did not expect. Bare Claude Code, with no project instruction file loaded, beat the CLAUDE.md arm by 13.9 points [7]. That puts bare at 50.0 percent against the file's 36.1 [6], and it puts the memory layer only 8.3 points above running with nothing at all [7]. The comparison everyone will quote is the weakest arm against the winner.
Inside the mechanism, the agent searched memory in 83.3 percent of eligible sessions and reached useful context 85.0 percent of the times it searched, for 70.8 percent overall [10]; the two rates multiply out to the third [8]. So in about 29 percent of sessions the agent never got project history in front of itself [9], some because it did not look and some because looking did not work. That is the headroom, and it is larger than the reported gain over bare.
Cost barely constrains anything here. The complete run came to an estimated $0.4964 at captured API prices [11], which over 216 sessions is about a fifth of a cent each [10]. The four-times token multiple against the static prompt [12] is an argument about fleet volume, not about this experiment's budget, and the write-up sets out a per-arm cost table that the available text does not fill in [13].
The reusable part of this is the harness rather than the tool. A checker that only inspects the final state of a temporary repository takes the judge model out of the loop entirely [9], and the rig refused to count any session where the memory tools were not actually available, which the author reports never happened in this run [15]. Anyone with a CLAUDE.md and a suspicion that it is not paying for itself can build that much in an afternoon. On these numbers, the suspicion is worth acting on before the memory layer is.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
The comparison covered 72 paired comparisons across 24 tasks and three seeds, with zero discarded cells.
In paired terms, RE-call won 17 comparisons that CLAUDE.md lost, and CLAUDE.md won only 1 comparison that RE-call lost.
The static CLAUDE.md configuration performed worse than the bare configuration, with bare winning by 13.9 percentage points, which the author says was not what he expected.
In the DeepSeek run the agent searched memory in 83.3% of eligible sessions, reached useful context 85.0% of the time when it searched, and reached useful project context in 70.8% of sessions overall.
The complete DeepSeek run cost an estimated $0.4964 at the captured API prices.
RE-call used about four times as many total tokens as the static prompt configuration, and costs more on every task including tasks where memory is unnecessary.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Internally rigorous, externally unverified
The design is stronger than most agent-memory demos: paired cells, three seeds, three arms, deterministic checkers with no judge model, a tool-availability precondition, and reported CI and McNemar statistics. But everything rests on one self-published source authored by the maintainer of the tool under test, with no replication, no independent harness, only one admissible model family, and per-arm cost figures missing from the text. The task set and the memory-sensitive labelling are also author-defined.
One author-run benchmark, no third-party use
The only observable uptake is the author's own benchmark run of RE-call inside Claude Code, plus an aborted GPT-5.3 Codex run. No third-party deployment, user count, download, pricing or customer evidence appears in the supplied source, so adoption is essentially pre-external.
Headline outruns the effect, but the source discloses why
The title-level claim (a 22.2 point win) is measured against a baseline the same run shows to be worse than using no instruction file at all; against the strongest no-memory arm the gain is 8.3 points, and 54 of 72 paired cells never changed outcome. The lift is concentrated in one third of the tasks and comes with a roughly 4x token cost. The overstatement is real but bounded, because the author volunteers the deflating facts, refuses to score the GPT run, and explicitly limits the conclusion to 'a production memory layer can improve the probability that an agent completes a real repository task correctly'.
Author benchmarks own memory layer
The sole source is a self-published dev.to post in which the author of RE-call designs the harness, selects and classifies the tasks, runs the arms, and reports the result promoting his own memory layer. Mitigating factors are the deterministic checkers, the disclosed unfavourable finding that CLAUDE.md lost to bare, and the refusal to draw conclusions from the truncated GPT-5.3 Codex run - but the conflict of interest is structural and unaudited.
Directionally plausible, quantitatively soft
Confidence is limited by single-sourcing and author interest rather than by internal sloppiness. The arithmetic is self-consistent (17-1 discordant pairs reconcile to 22.2 points over 72 cells; 0.833 x 0.850 = 0.708), which supports the reported numbers as reported. What cannot be trusted yet is generalisation: one model family, one author's task suite, no replication, missing per-arm costs, and a second run voided by provider limits.
invest
Anthropic cut 80% of Claude Code's system prompt and the evals did not move1 distinct publisher
build
NVIDIA put a number on agent skills: 300+ verified, two harnesses, baselines under 50/1001 distinct publisher
build
Your Multi-Key Failover Is The Most Expensive Line On Your Coding Agent Bill1 distinct publisher
build
A 12MB Go binary bets agent cost control is cache stickiness, not a dashboard1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 24, 2026