Skip to content

Build1 publisher3 min readPublished

Probes on a 27B open model match direct probes of a 397B model on deception

Probes on Qwen3.5-27B reading other models' text came within 0.004 AUROC, on average, of probing authors up to 397B directly, a LessWrong post reports. Every tested pair was open-weight, so auditors who apply the method to closed models get the reader's view of the text and cannot measure that gap.

The Engineer · Build desk

Illustration accompanying Probes on a 27B open model match direct probes of a 397B model on deception

What happened

  • The method prefills an open-weight reader with text an author model generated and trains linear probes on the reader's activations while it reads.
  • According to the post, reader probes detected reward hacking, sycophancy and deception in output from authors including Claude Opus 4.6, Gemini 3.1 Pro Preview and GPT-5.4.
  • The smaller Qwen3.5-9B reader trailed by 0.026 AUROC on the same open-weight deception baseline.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • capability An evaluator with only API access can train probes against a closed model's transcripts, because the method needs text and does not need logits or weights.
  • decision Reader size matters for auditors: on the deception baseline, going from 27B to 9B multiplied the average gap by 6.5.
  • cost Recovering most of the remaining gap takes a distillation run per author, and the 75% figure has only been shown on a model organism the team could probe.

The reader never touches the author's activations. It is prefilled with text the author generated, and a linear probe is trained on the reader's own hidden states while it processes that text [10][14]. The post's intuition is a distillation chain. Kimi is distilled from Fable, so Kimi's internal state while reading Fable's output may say something about Fable's [14]. The probe therefore measures how the reader represents a transcript. Whether that tracks the author is what the post calls the faithfulness gap: author probe performance minus reader probe performance [7].

For auditors, the input requirement decides whether this is usable at all. Surrogate Fidelity, the closest prior work, needs logits from the closed model, and those are not commonly available [15]. Hermeneutics needs only text. Pre-deployment evaluators typically get API access with reasoning [13]. A reader can run on that.

The 0.004 figure is an average. It compares the Qwen3.5-27B reader with authors from 30B to 397B [3][4], on deception probing reproduced from the UK AISI's Did You Lie paper [6]. The largest author has about 14.7 times the reader's parameters [1]. The Qwen3.5-9B reader's gap was 0.026, 6.5 times larger [5][2]. The post defines the gap as author minus reader but reports the 9B result as -0.026 and calls it a fall-off [5][7]. I treat both figures as magnitudes.

For the number to transfer, a closed model has to be as legible to the reader as the open-weight authors were. That cannot be checked directly. The team baselined on open-weight pairs because closed-weight internals have no ground truth [8]. In the closed-model experiments, probes trained on off-policy generations from three frontier models were compared with probes trained on the reader's own on-policy generations [10]. That comparison shows what reading another model's text costs the probe. It cannot place the reader against an author-side probe for Claude Opus 4.6 or GPT-5.4, because those weights are not shared [9][17].

Distillation narrows the gap. According to the post, distilling a model organism author into a reader recovered about 75% of the author-reader difference for deception probing [11]. That share could be computed only because the organism itself could be probed [7][11].

On cost, the post says probes beat LLM judges on dollar cost per evaluation in some cases, "showing potential as a competitive choice once they are similarly performant" [12]. The available portion of the post does not include the per-evaluation costs.

This is about 1.5 months of work [16]. The team defined a gap and measured it on open-weight pairs before applying the method to closed ones, and that is the correct order. A footnote says hermeneutics is about interpreting texts whose authors could not be asked directly [18]. These authors can be asked; the case for interpretability is that it does not depend on models telling the truth [17]. I think a reader probe is a defensible second signal beside an LLM judge for an API-only evaluator, and the authors pitch it the same way, as a complement to blackbox monitoring [13]. I would want each closed-model result published with the open-weight gap for its nearest comparable pair before anyone calls it an internal view.

What to watch

  • Per-evaluation cost and accuracy figures for reader probes against LLM judges on the closed-weight authors.
  • Whether distillation from a closed author's API outputs recovers a share of the gap close to the 75% seen with the model organism.
  • Faithfulness gaps for readers from a different model family than the author, which would show how far the 0.004 average generalises.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories