Build1 distinct publisher3 min readPublished
A published MCP memory server makes the model propose facts and lets prompt-free code decide what is stored. The author's near-miss the night before launch is the part worth reading.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
The gate that refuses "under duress" is the cheap win [5]. Any comparison strict enough to reject an added clause will reject that one. The expensive part of this design sits on the other side of the interaction, in what the store says when it holds nothing, and that is where the author says he was nearly caught out.
He calls the post the honest version and points at a night before publication when the thing lied to him [14]. The concrete fix he attaches to that stretch is the refusal redirect: an abstention that names the grounded terms it does hold and the next action worth taking, instead of a bare no [11][13]. His reasoning is that most callers are agents, and an agent cannot browse the store to work out a better question, so a flat refusal is where the exchange dies [12].
Look at what the repaired path is doing. The abstention reports that two claims about Dana Kim exist and lists the terms they ground on [11]. That is an assertion about the contents of the store, and unlike an admitted claim it does not arrive with a document hash and a byte range [8]. The admit path is audited. The refusal path is prose about the audit. My reading is that if a system of this shape tells its user something the user cannot check, it will be there.
The verify output hints at the same seam. Success prints one of one receipts checkable and verified; the failure case prints zero of one checkable, then FAILED [9]. "Checkable" is carrying weight. A receipt whose source document has gone missing is neither verified nor caught, and that counter is the only place the distinction surfaces.
His diagnosis, that the component doing the hallucinating is also the one keeping the record [17], is what puts the check at write time rather than in a read-time reranker or LLM judge [7]. The security claim follows from placement: there is no model and no prompt on the deciding side, so injection and jailbreaks find nothing to persuade [3][6].
That leaves the write contract as the thing that decides whether anyone else can use it. A claim has to arrive with the exact text it grounds on [4], and the receipt is a document hash plus a start and end offset [8], 71 bytes wide in the published example [1]. A retrieval stack that hands the model reranked paraphrase has nothing to give that function. What it buys is provenance that can go red on purpose, rather than a chunk pointer that keeps looking fine after the underlying text drifts [9][10].
Ranked by verification strength, evidence, and original report placement.
The author published an MCP memory server in which the model does not decide what is remembered; it proposes, and deterministic code decides what is admitted.
The stated rule is one sentence: the model proposes, deterministic code decides, and nothing ungrounded is committed.
The deciding code contains no model and no prompt anywhere in it, according to the author.
The agent does not simply assert a fact: it asserts a claim and quotes the exact text it is grounding on, and a gate checks whether the quote actually supports the claim.
A worked example shows remember(claim="Priya joined Acme in 2019 under duress.", evidence="Priya Raman joined Acme in 2019 as a logistics analyst.") returning REFUSED (asserts_more_than_evidence), because "under duress" is not in the evidence.
An admitted claim is stored with a receipt consisting of (document_hash, byte_start, byte_end); the published example shows bytes [0:71] of a sha256 document hash for the claim "Dana Kim has a cat named Pepper."
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Author-demonstrated only
All substantiation is first-person: design description, tool transcripts and the author's own pre-launch testing, in one self-published post. Mechanism claims (byte-range receipts, deterministic refusal, abstention redirect) are illustrated with concrete outputs, which is better than bare assertion, but nothing is independently reproduced, audited or benchmarked, and the broader security and RAG-comparison arguments have no supporting measurement at all.
Just published, no usage signal
The only adoption fact in the material is publication itself: an MCP server released with a package two days old on PyPI. No installs, downloads, external users, integrations or production deployments are disclosed, and the author's own clean-environment test surfaced a silent write failure days before publication, so there is no evidence of use beyond the author's machine.
Strong guarantees, thin verification
Language such as provenance you can falsify, cheap and permanent, and the assertion that prompt injection and jailbreaks give an attacker no purchase, sets a guarantee-level bar that a two-day-old package with no external validation cannot yet clear. The gap is real but moderated by unusual self-disclosure: the author volunteers the code's age and the ADMITTED-while-storing-nothing bug, which keeps the piece well short of pure promotion.
Self-promotional but disclosed
The source is the author writing about a package he just published, so the piece functions as launch material for his own project and carries a clear interest in the design being seen as sound. The incentive is fully visible rather than hidden -- authorship is first-person throughout, and the post spends its closing section on the project's weaknesses, including a defect found the night before publication, which is atypical of promotional writing. No commercial terms, employer, funding or sponsorship are disclosed either way.
Low - single self-reported account
Confidence is limited by structure rather than by internal inconsistency: one publisher, one author, no corroboration, and every mechanism claim resting on transcripts only the author can reproduce. What the source is being trusted for -- what the software does and how new it is -- is exactly what a builder is best placed to report, and the self-critical closing section raises credibility, so the descriptive claims are reasonably reliable while the guarantee and comparison claims are not yet assessable.
build
255 tool schemas, 91K tokens: pricing the two MCP costs nobody budgets1 distinct publisher
build
Rate limit your MCP servers, because a retrying agent turns one error into a billing incident1 distinct publisher
build
Thirty minutes a day, and none of it from letting the agent write Swift1 distinct publisher
build
The MCP transport your search results teach has been deprecated since March1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 24, 2026