Build1 publisher3 min readPublished
Probes on a 27B open model match direct probes of a 397B model on deception
Probes on Qwen3.5-27B reading other models' text came within 0.004 AUROC, on average, of probing authors up to 397B directly, a LessWrong post reports. Every tested pair was open-weight, so auditors who apply the method to closed models get the reader's view of the text and cannot measure that gap.
The Engineer · Build desk

What happened
- The method prefills an open-weight reader with text an author model generated and trains linear probes on the reader's activations while it reads.
- According to the post, reader probes detected reward hacking, sycophancy and deception in output from authors including Claude Opus 4.6, Gemini 3.1 Pro Preview and GPT-5.4.
- The smaller Qwen3.5-9B reader trailed by 0.026 AUROC on the same open-weight deception baseline.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- capability An evaluator with only API access can train probes against a closed model's transcripts, because the method needs text and does not need logits or weights.
- decision Reader size matters for auditors: on the deception baseline, going from 27B to 9B multiplied the average gap by 6.5.
- cost Recovering most of the remaining gap takes a distillation run per author, and the 75% figure has only been shown on a model organism the team could probe.
The reader never touches the author's activations. It is prefilled with text the author generated, and a linear probe is trained on the reader's own hidden states while it processes that text [10][14]. The post's intuition is a distillation chain. Kimi is distilled from Fable, so Kimi's internal state while reading Fable's output may say something about Fable's [14]. The probe therefore measures how the reader represents a transcript. Whether that tracks the author is what the post calls the faithfulness gap: author probe performance minus reader probe performance [7].
For auditors, the input requirement decides whether this is usable at all. Surrogate Fidelity, the closest prior work, needs logits from the closed model, and those are not commonly available [15]. Hermeneutics needs only text. Pre-deployment evaluators typically get API access with reasoning [13]. A reader can run on that.
The 0.004 figure is an average. It compares the Qwen3.5-27B reader with authors from 30B to 397B [3][4], on deception probing reproduced from the UK AISI's Did You Lie paper [6]. The largest author has about 14.7 times the reader's parameters [1]. The Qwen3.5-9B reader's gap was 0.026, 6.5 times larger [5][2]. The post defines the gap as author minus reader but reports the 9B result as -0.026 and calls it a fall-off [5][7]. I treat both figures as magnitudes.
For the number to transfer, a closed model has to be as legible to the reader as the open-weight authors were. That cannot be checked directly. The team baselined on open-weight pairs because closed-weight internals have no ground truth [8]. In the closed-model experiments, probes trained on off-policy generations from three frontier models were compared with probes trained on the reader's own on-policy generations [10]. That comparison shows what reading another model's text costs the probe. It cannot place the reader against an author-side probe for Claude Opus 4.6 or GPT-5.4, because those weights are not shared [9][17].
Distillation narrows the gap. According to the post, distilling a model organism author into a reader recovered about 75% of the author-reader difference for deception probing [11]. That share could be computed only because the organism itself could be probed [7][11].
On cost, the post says probes beat LLM judges on dollar cost per evaluation in some cases, "showing potential as a competitive choice once they are similarly performant" [12]. The available portion of the post does not include the per-evaluation costs.
This is about 1.5 months of work [16]. The team defined a gap and measured it on open-weight pairs before applying the method to closed ones, and that is the correct order. A footnote says hermeneutics is about interpreting texts whose authors could not be asked directly [18]. These authors can be asked; the case for interpretability is that it does not depend on models telling the truth [17]. I think a reader probe is a defensible second signal beside an LLM judge for an API-only evaluator, and the authors pitch it the same way, as a complement to blackbox monitoring [13]. I would want each closed-model result published with the open-weight gap for its nearest comparable pair before anyone calls it an internal view.
What to watch
- Per-evaluation cost and accuracy figures for reader probes against LLM judges on the closed-weight authors.
- Whether distillation from a closed author's API outputs recovers a share of the gap close to the 75% seen with the model organism.
- Faithfulness gaps for readers from a different model family than the author, which would show how far the 0.004 average generalises.