Skip to content

Build1 publisherNot yet confirmed elsewhere3 min readPublished

Filtering noisy gradients lifts the J++ Lens to 55% recall on models' hidden variables

Anthropic researchers' J++ Lens puts the right hidden intermediate variable in its top 10 readouts 55.2% of the time, against 35.7% for the J-Lens. The authors pitch it for monitoring eval awareness, but the score comes from prompts with known answers.

The Engineer · Build desk

How we use AISend a correction

Photograph accompanying Filtering noisy gradients lifts the J++ Lens to 55% recall on models' hidden variables
Photo: lesswrong.com

What happened

  • The original J-Lens gave coherent readouts only in the latter half of a model's layers and failed to surface many intermediate variables the authors expected.
  • A second baseline, the R-Lens, reached 37.7% on the same recall@10 measure.
  • The authors report the improvement across language models ranging from 9 billion to 284 billion parameters.
  • In one worked example, J++ surfaced the intermediate variable at layer 24, where the J-Lens needed layer 40.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • cost Switching costs existing J-Lens users little engineering, since the authors describe J++ as a drop-in replacement and ship code, a demo and open-source lenses.
  • constraint The right intermediate is still missing from the top ten readouts 44.8% of the time, so a readout that shows nothing suspicious is weak evidence that nothing is there.
  • capability If the layer-24 result holds beyond the authors' example, monitors could inspect the first half of a model's layers, where J-Lens readouts were incoherent.

The J-Lens, introduced by Gurnee et al. (2026), returns the tokens a model is most disposed to verbalise given a single activation [3]. It gets there through a Jacobian averaged over many activations [4]. Kola Ayonrinde of Anthropic Fellows and Jack Lindsey of Anthropic found that noisy gradients pile up in that average and corrupt it [c17, c6]. Their fix has three parts: Jacobians from activations that give poor readouts are down-weighted, Layer-wise Relevance Propagation is applied to the backward pass when the Jacobians are computed, and non-semantic tokens are removed from the readout [7]. Only the first two act on the Jacobian. The third filters the token list the lens returns, so the gradient-filtering label also covers a cleanup of the output [7].

The order of the work is the part I would copy. Before touching the tool, the authors set out four reasons the J-Lens might miss an intermediate variable: the workspace hypothesis is a poor model of language model cognition, the workspace lacks a linear representation of the variable, the model does not represent it at all, or the lens is imperfect [8]. They then showed that a substantial fraction of intermediate variables can be extracted linearly [9]. That evidence points at the lens. Next they showed that filtering out the Jacobians from a subset of activations raises lens performance [10].

The workspace is worth reading because Gurnee et al. tied it to behaviour. Given the prompt "Fact: The number of legs on the animal that spins webs is", the J-Lens surfaced "spider" and "legs" at intermediate layers before the model predicted 8 [11]. Patching in the J-Lens vector for "ant" in place of "spider" flipped the answer to 6, the right count for an ant [11].

On the authors' extraction tasks, J++ beats the J-Lens by 19.5 points and the R-Lens by 17.5 [d1, d2]. The metric is recall@10, so the correct intermediate has to appear among the top ten readout tokens [1]. The targets are known in advance. In one example the intended path runs from the prompt to "heart" to "four chambers" [17]. A monitor watching for eval awareness has no known target to check against. For the score to carry over, the concept would have to sit in the workspace the way a multi-hop intermediate does. Someone would also have to spot it in a ten-token list without knowing what to look for.

Ayonrinde and Lindsey wrote, "We believe that more faithful lenses can be useful in monitoring for eval awareness and unverbalised scheming." [21] The results in the post come from the latent variable extraction tasks [1]. The post does not say how a poor readout is scored for down-weighting, or how the 55.2% splits across the three changes and the range of model sizes [c7, c14].

What to watch

  • An ablation splitting the 19.5-point gain between the two Jacobian changes and the removal of non-semantic tokens.
  • A test of J++ on eval-awareness or scheming cases where the concept to find is not specified in advance.
  • Per-model results showing whether the gain holds at the 284B end of the range.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence45
Adoption
Insufficient
Hype gap+20
Incentives40
Confidence50
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    The J++ Lens surfaces the correct intermediate readouts at 55.2% recall@10 on latent variable extraction tasks, compared with 35.7% for the J-Lens.

    ReportedSupportedSource: Ayonrinde and Lindsey, J++ Lens paper postView cited source
  2. [2]

    The R-Lens reached 37.7% recall@10 on the same intermediate readout measure.

    ReportedSupportedSource: Ayonrinde and LindseyView cited source
  3. [3]

    The J-Lens, introduced by Gurnee et al. (2026), extracts the tokens that the model is most disposed to verbalise, conditional on a given activation.

    ReportedSupportedView cited source

Sources

1 independent publisher whose own reporting we read for this story.

  1. lesswrong.com

    1 article · October 8, 2026

    [Paper] J++ Lens: Jacobian Filtering Enables More Faithful Workspace Lenses

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Entities

Loading related stories