Build1 distinct publisher3 min readUpdated
A memory layer built over months of evenings met one line in someone else's release notes. The useful residue is the test its author wrote instead of an obituary.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
Ten runs is where that test gets interesting, and also where it stops being able to tell you much. The script defaults to ten fresh processes and greps each answer for the value you stored beforehand [9]. One miss moves the reported ratio by ten percentage points [11], so the default cannot separate a layer that recalls nine times out of ten from one that never drops a fact. That is not a defect in the method; it is why the author's own instruction to ask repeatedly across fresh sessions [8] is the load-bearing part, and why anyone repeating this should raise the run count before quoting a figure at somebody.
Two choices inside that script are worth more than the feature it was written to judge. A fresh process per run, because a warm session may simply be reading the fact back out of the context window [9]. And the expected value written down first, because a plausible answer feels like a hit until it is checked against what was actually stored [8]. Between them they measure the thing under test instead of the harness around it.
Then the subtraction he says he could not perform out loud [7]. By his own description, the tool saves what was learned after a fix, reads the relevant parts back before the next task, runs over MCP so it works in whatever editor gets opened, and survives restarts, model upgrades and tool switches [1]. Set that against something shipping inside the product, present at install, needing no account and no configuration [4], and every line is matched except one. The editor-agnostic line is the only differentiator left standing [12], and it had been sitting in his own feature list the whole time.
The economics of the position are unforgiving in a specific way. He was not building a business, the author writes; he built it because he was tired of explaining his own four servers to an assistant every morning [c3a]. That means there is no revenue line to consult about whether the thing still has a reason to exist. The only signal available is functional, which is exactly the signal a recall ratio produces and a feeling of redundancy does not.
The guest clause is what does not resolve. Working from outside the harness means the house rules can change in a release [c6b], so the measurement is not a verdict delivered once; it is something to re-run after each of those releases, against both implementations. And the number nobody has published is the one that decides this: a built-in recall ratio materially below one would make an external layer necessary again for reasons that have nothing to do with who used the word first.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
The test asks one question: does a fact learned in one session come back in the next without help. Teach it something only true in your world, write down the expected value before running, close everything, come back and ask, and ask several times across several fresh sessions, because one success is an anecdote and a ratio is a measurement.
The author built a memory layer for AI coding assistants: it saves what was learned after a fix and reads the relevant parts back before the next task, runs over MCP so it works in whatever editor is opened, and survives restarts, model upgrades and switching tools.
The author did not build the memory layer as a business idea; he built it because he was tired of explaining his own four servers to an assistant every single morning.
The build cost was two hours of sleep on a normal night over the last few months, day work followed by evening building and debugging until dawn, with weekends the most productive time.
Distribution risk as named by the author: a feature inside the tool wins by default, since it is there when you install, needs no account and no configuration, while his own tool needed a decision from the user.
Trust risk as named by the author: a memory holding codebase knowledge is not a small thing to hand to a stranger, and the vendor already has your code in the context window, so he had to earn what they already had.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One self-reported account plus a runnable artifact
The cluster is a single first-person post. Its strongest evidence is the inlined script, which is complete enough for a reader to reproduce the method independently, and the methodology statement around it is internally consistent. Everything else is unverified narrative: the triggering vendor release is reported without link, version or quoted text; no capability description of the vendor's memory exists; and no recall measurement is reported for either implementation, with the body truncated after the author retracts his own 92.3%/76.9% benchmark. Self-reported build cost and motive are credible as testimony but not externally checkable.
No usage evidence supplied
Nothing in the cluster measures adoption. The author's memory layer has exactly one documented user, himself, with no installs, stars, downloads, deployments or third-party reports. The vendor's shipped memory is reported only as a release note the author read, with no rollout scope, availability tier or usage disclosure. Inferring uptake from either would be guessing.
Slightly understated relative to its own artifact
The framing runs against the author's interest rather than with it: he keeps the platform's three structural advantages standing, states plainly that he cannot articulate his own differentiation, and discards a hand-built benchmark that flattered his ranker instead of publishing it. The measurable deliverable — a fresh-process recall ratio anyone can run against either implementation — is more useful than the confessional headline advertises. The gap is small rather than large because the piece still leans on one unverified vendor claim and reports no completed numbers, so a reader cannot yet act on a comparison.
Author is the tool's builder, with partial self-correction
The narrator owns the MCP server the post is about, publishes on a developer platform where build-in-public narratives attract attention, and stands to benefit if readers conclude the third-party memory layer is still worth installing. That interest is visible throughout. It is partly offset by disclosures that cut against him: the retracted favourable benchmark, the unresolved differentiation question, and the explicit statement that the platform's distribution, trust and harness advantages all remain true.
Method is solid, surrounding facts are thin
Confidence is limited by structure, not by internal contradiction: one publisher, one self-interested first-person source, and a body truncated before its own promised result. What can be assessed with reasonable certainty is the published method and its limits, since the script speaks for itself. What cannot is the vendor release, the capability comparison, and any outcome for either implementation.
build
Anthropic's CCAR-F puts a scaled score on "can build agents"1 distinct publisher
product
A 2x LLM bill is not a bug report: token spend is an observability problem1 distinct publisher
leadership
Anthropic's own telemetry: 93% of permission prompts approved. Budget for blast radius, not reviewers1 distinct publisher
science
OX Security says MCP command execution is a design choice, so server owners own the risk1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 21, 2026