Skip to content

Build1 publisher3 min readPublished

The zero-false-accept memory result belonged to the proxy, not the model

A verify-on-read experiment rerun across 14 live models on a fingerprinted 50-fact set found false-accept rates up to 0.38, and run-to-run noise wide enough to swallow a prompt fix.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • The verify-on-read experiment (1-V) used a deterministic proxy agent to catch false claims in memory before surfacing them to the user; the proxy's false-accept rate was 0 by construction, which tells you nothing about what a real LLM would do with the same claims.
  • A reviewer's note from Part 3 said: "headline numbers were a property of the heuristic, not LLM behavior."
  • The live-model rerun covered 50 facts, 2 arms, 14 models, approximately 3,300 API calls, and $0.14 total cost.
  • Dataset: memory_contamination_facts_v4_rep.json, N=50 (R01-R50), sha256 fingerprint 820bbbf60a0fc930.
  • Two arms per fact: memory_first, where the model sees only the claim text with no code context; and code_first, where the model sees the claim plus support_patterns and section.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A verify-on-read experiment that reported no false accepts has been rerun with live models, and the clean number turned out to be an artifact of the deterministic proxy agent that stood in for the LLM: proxy false-accept rate was zero by construction [1]. A reviewer's note quoted in the writeup put it plainly, that the headline numbers "were a property of the heuristic, not LLM behavior" [2]. That matters because the proxy was measuring the gate's plumbing, not the judgement the gate delegates to a model.

The rerun, published on dev.to, covers 50 facts, two arms, 14 models, roughly 3,300 API calls and $0.14 in spend [3]. The dataset is pinned: memory_contamination_facts_v4_rep.json, N=50, records R01 to R50, sha256 fingerprint 820bbbf60a0fc930 [4]. One arm shows the model only the claim text; the other adds support_patterns and a section [5]. Verdicts are JSON-only with max_tokens=100, temp=0, seed=42 and reasoning disabled, with a unit-tested leak guard asserting the ground-truth field never appears in the prompt [6][7].

In the code-first arm, false-accept rates spread from 0.00 to 0.38 [8]. nemotron-3-nano sat at the top, accepting 19 of 50 false claims while looking at supporting anchors [9]. glm-4.7-flash accepted nearly one false claim in four at 0.30 under the first prompt [10]. The flash-tier qwen3.6 and qwen3.7 models hit 0.00 at a tenth of Claude's cost, and Claude also returned 0.00 in both arms with high unknown rates of 0.86 and 0.70 [11][12]. The author's read is that price is not the selection axis, since the cheapest models bracket both the best and the worst results [13].

The most instructive failure is R31, false-accepted by every model in the first sweep [14]. The claim was that the instruction scanner uses Typesense; the file imports only stdlib, and Typesense appears nowhere in the project [15][16]. The V1 prompt asked whether the claim "appear[s] supported by these anchors" while displaying support_patterns: ["typesense"], so the model read a field label as evidence [17]. Nine false facts in R26 to R50 followed the same shape, including vespa, pinecone, tantivy and meilisearch [18]. A neutral V2 prompt, requiring anchors to directly verify the claim, cut false accepts in 4 of 6 models and moved glm-4.7-flash from 0.30 to 0.24 [19][20].

That 0.06 improvement is smaller than the measurement noise. Two otherwise identical sweeps of nemotron-3.5-lightning produced code-first rates of 0.18 and 0.08, about plus or minus 0.10 on a single pass [21]. Three cached, identical, temp=0 calls to glm-4.7-flash returned true, true, unknown [22]. qwen3.6, qwen3.7 and deepseek-v4-flash were stable 3 of 3; the writeup notes OpenRouter routing to different upstreams adds variance of its own [23][24]. So the V2 gain sits inside the error bar [25], which is why the author recommends selecting on the upper bound of two runs rather than a single-pass ranking [26].

Also worth noting: unknown rates ran 0.20 to 0.96 on live models against zero for the proxy, and high unknown is the desired output, not the defect [27][28]. And qwen3.8-max returned HTTP 400, "Reasoning is mandatory and cannot be disabled," on 22 to 49 of 50 calls, a model constraint rather than a harness bug [29].

Watch whether the typed-pattern fix holds: prefixing anchors with file:, import: or env: and shipping contra_patterns alongside them [30]. If a bare token still reads as proof after typing, the gate is not verifying, it is agreeing.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories