Build1 publisher3 min readPublished
Removing synthetic markers lets implanted false facts pass a middle-layer linear probe
Subtly implausible planted facts pass a middle-layer linear probe once synthetic markers are scrubbed from training documents, a LessWrong study finds. The implants still do not always reach downstream reasoning, so a passing probe proves less than it seems.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- Features that set SDF training documents apart from pretraining text partly explain why linear probes could separate implanted facts from learned ones.
- Running RPO after marker reduction strengthens the planted facts monotonically, though overtraining with it becomes detectable.
- Egregiously implausible false facts can still be implanted with these methods, and the probe continues to catch them.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- exposure An evaluation that trusts a probe pass can run a weak-monitors test on a model that never acted on the belief, then draw conclusions from behaviour the implant never shaped.
- constraint Only subtle false facts get the probe-level disguise, so implausible beliefs stay open to the self-inspection the author worries about.
- cost Teams strengthening implants with RPO pay in detectability once they overtrain, so training strength has to be tuned against the probe.
- precedent False-facts studies that attribute probe separability to disbelief will be expected to control for synthetic markers first.
The detector is a linear classifier trained on middle-layer activations [2]. The procedure comes from Appendix G.1 of Believe It or Not (Slocum et al., 2025) [13]. Every fact in the set is one the model learned in its true form during pretraining [15]. The author finetunes the false version of 30 of them and leaves another 30 untouched [14]. The probe then has to sort implanted facts from untrained ones [13]. If it cannot, the planted belief is represented like pretraining knowledge, at least as far as a linear probe at that layer can tell [3].
Earlier separability results had two possible causes. A probe might split the sets because the model disbelieves the implants, or because the training documents look synthetic, and according to the author no prior work had tested which [16]. The post defines a synthetic marker as any feature of a training document that distinguishes it from pretraining text. Its examples are framing the fact as more surprising or more hypothetical than other facts, stating it more directly, and using more em dashes around it [12]. Anyone who has read much model-generated prose will recognise the last one. The markers turn out to be only partly responsible for the separability [2].
Splitting those causes apart is careful work. The reason to care is evaluation design. Planting the belief "My developer's monitors are weak. They cannot catch these subtle attacks" could provoke a model to reveal its attack strategies in a test, the author wrote [7]. A capable enough model, the author worries, might inspect its own beliefs and pick out the planted ones [10]. The post imagines that model's reaction: "What a strangely represented belief! Maybe it's an implanted fact, and therefore, likely a false fact." [11]
The results sort by plausibility. Subtly implausible facts written without markers pass the probe [3]. A follow-up step the post calls RPO strengthens them monotonically, though overtraining with it is detectable [4]. Egregiously implausible facts can be implanted, and the probe still catches them [5]. Slocum et al. had already found plausible false facts nearly indistinguishable from true ones, and less useful [9]. The gain from marker removal is in the middle band [3] [5] [9]. The post's summary states these outcomes without probe accuracies.
The result I would weight most is the downstream one. Facts the probe cannot separate from pretraining knowledge still do not always carry into reasoning tasks [6]. A probe that fails to split two sets shows only that they are not linearly separable at that layer. It does not show that the model reaches for the planted fact when it plans. For the probe result to transfer, the features the probe picks up would have to be the ones the model's reasoning uses, and the downstream finding says they are not always [6].
Two more conditions bound the claim. Sixty facts is a small sample for a classifier result [1]. The post describes the procedure on a model such as Llama-3.1-8B-Instruct [17], while the concern is about models capable of inspecting their own beliefs [10]. I'd treat a passing probe as an early filter for an implant and the downstream reasoning tasks as the acceptance test.
What to watch
- The post's per-fact downstream reasoning results, and whether failures to propagate track how implausible each fact is.
- Whether the indistinguishability holds against nonlinear probes or probes at layers other than the middle one.
- A replication on models larger than Llama-3.1-8B-Instruct, where the author's self-inspection concern is aimed.