Build1 publisherNot yet confirmed elsewhere3 min readPublished
Filtering noisy gradients lifts the J++ Lens to 55% recall on models' hidden variables
Anthropic researchers' J++ Lens puts the right hidden intermediate variable in its top 10 readouts 55.2% of the time, against 35.7% for the J-Lens. The authors pitch it for monitoring eval awareness, but the score comes from prompts with known answers.
The Engineer · Build desk
What happened
- The original J-Lens gave coherent readouts only in the latter half of a model's layers and failed to surface many intermediate variables the authors expected.
- A second baseline, the R-Lens, reached 37.7% on the same recall@10 measure.
- The authors report the improvement across language models ranging from 9 billion to 284 billion parameters.
- In one worked example, J++ surfaced the intermediate variable at layer 24, where the J-Lens needed layer 40.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- cost Switching costs existing J-Lens users little engineering, since the authors describe J++ as a drop-in replacement and ship code, a demo and open-source lenses.
- constraint The right intermediate is still missing from the top ten readouts 44.8% of the time, so a readout that shows nothing suspicious is weak evidence that nothing is there.
- capability If the layer-24 result holds beyond the authors' example, monitors could inspect the first half of a model's layers, where J-Lens readouts were incoherent.
The J-Lens, introduced by Gurnee et al. (2026), returns the tokens a model is most disposed to verbalise given a single activation [3]. It gets there through a Jacobian averaged over many activations [4]. Kola Ayonrinde of Anthropic Fellows and Jack Lindsey of Anthropic found that noisy gradients pile up in that average and corrupt it [c17, c6]. Their fix has three parts: Jacobians from activations that give poor readouts are down-weighted, Layer-wise Relevance Propagation is applied to the backward pass when the Jacobians are computed, and non-semantic tokens are removed from the readout [7]. Only the first two act on the Jacobian. The third filters the token list the lens returns, so the gradient-filtering label also covers a cleanup of the output [7].
The order of the work is the part I would copy. Before touching the tool, the authors set out four reasons the J-Lens might miss an intermediate variable: the workspace hypothesis is a poor model of language model cognition, the workspace lacks a linear representation of the variable, the model does not represent it at all, or the lens is imperfect [8]. They then showed that a substantial fraction of intermediate variables can be extracted linearly [9]. That evidence points at the lens. Next they showed that filtering out the Jacobians from a subset of activations raises lens performance [10].
The workspace is worth reading because Gurnee et al. tied it to behaviour. Given the prompt "Fact: The number of legs on the animal that spins webs is", the J-Lens surfaced "spider" and "legs" at intermediate layers before the model predicted 8 [11]. Patching in the J-Lens vector for "ant" in place of "spider" flipped the answer to 6, the right count for an ant [11].
On the authors' extraction tasks, J++ beats the J-Lens by 19.5 points and the R-Lens by 17.5 [d1, d2]. The metric is recall@10, so the correct intermediate has to appear among the top ten readout tokens [1]. The targets are known in advance. In one example the intended path runs from the prompt to "heart" to "four chambers" [17]. A monitor watching for eval awareness has no known target to check against. For the score to carry over, the concept would have to sit in the workspace the way a multi-hop intermediate does. Someone would also have to spot it in a ten-token list without knowing what to look for.
Ayonrinde and Lindsey wrote, "We believe that more faithful lenses can be useful in monitoring for eval awareness and unverbalised scheming." [21] The results in the post come from the latent variable extraction tasks [1]. The post does not say how a poor readout is scored for down-weighting, or how the 55.2% splits across the three changes and the range of model sizes [c7, c14].
What to watch
- An ablation splitting the 19.5-point gain between the two Jacobian changes and the removal of non-semantic tokens.
- A test of J++ on eval-awareness or scheming cases where the concept to find is not specified in advance.
- Per-model results showing whether the gain holds at the 284B end of the range.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+20
- Incentives40
- Confidence50
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
The J++ Lens surfaces the correct intermediate readouts at 55.2% recall@10 on latent variable extraction tasks, compared with 35.7% for the J-Lens.
- [2]
The R-Lens reached 37.7% recall@10 on the same intermediate readout measure.
- [3]
The J-Lens, introduced by Gurnee et al. (2026), extracts the tokens that the model is most disposed to verbalise, conditional on a given activation.
- [4]
The J++ Lens filters noisy gradients before they enter the averaged Jacobian.
- [5]
The J-Lens tends to give coherent readouts only in the latter half of layers, and many intermediate variables expected in the workspace are not surfaced by it at all.
- [6]
The authors show that the J-Lens accumulates noisy gradients, corrupting the averaged Jacobian.
- [7]
The J++ Lens adds three changes to the J-Lens: down-weighting the Jacobians from activations that result in poor readouts; applying Layer-wise Relevance Propagation to the backward pass when computing Jacobians; and removing non-semantic tokens from the readout.
- [8]
The authors list four possible explanations for J-Lens failures: the hypothesised workspace is a poor model of language model cognition; the workspace lacks a linear representation of the variables; the model does not represent them at all; or the J-Lens is an imperfect tool.
- [9]
The authors provide evidence that a substantial fraction of intermediate variables can be linearly extracted.
- [10]
The authors show that filtering out Jacobians from a subset of the activations increases lens performance.
- [11]
Per Gurnee et al. (2026), for the prompt 'Fact: The number of legs on the animal that spins webs is', the J-Lens surfaces 'spider' and 'legs' in intermediate layers before the model predicts '8'; patching the J-Lens vector 'ant' in place of 'spider' flips the response from 8 to 6.
- [12]
In an example, the J++ Lens surfaces semantically relevant intermediate latent variables at layer 24, compared with layer 40 for the J-Lens and R-Lens.
- [13]
The J++ Lens enables more faithful readouts across models from 9B to 284B parameters.
- [15]
The authors release code, an interactive demo and open-source lenses.
- [16]
The paper's authors are Kola Ayonrinde of Anthropic Fellows and Jack Lindsey of Anthropic.
- [17]
In the authors' worked example, the intended journey hops from the prompt to 'heart' to 'four chambers'.
- [18]
The J++ Lens beats the J-Lens by 19.5 percentage points on recall@10.
- [19]
The J++ Lens beats the R-Lens by 17.5 percentage points on recall@10.
- [20]
At 55.2% recall@10, the correct intermediate is absent from the top ten readouts 44.8% of the time.
- [21]
"We believe that more faithful lenses can be useful in monitoring for eval awareness and unverbalised scheming."
ReportedInsufficientSource: Kola Ayonrinde and Jack Lindsey, TL;DR of the paper postView cited source
Sources
1 independent publisher whose own reporting we read for this story.
- lesswrong.com[Paper] J++ Lens: Jacobian Filtering Enables More Faithful Workspace Lenses
1 article · October 8, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.
Topics
- AI safety monitoringFollow
- Mechanistic interpretabilityFollow