Skip to content

Topic

Mechanistic interpretability

A research field that reverse-engineers neural networks' internal computations—circuits, features, representations—to explain how models produce outputs.

Current stories

build1 publisher

Probes on a 27B open model match direct probes of a 397B model on deception

Probes on Qwen3.5-27B reading other models' text came within 0.004 AUROC, on average, of probing authors up to 397B directly, a LessWrong post reports. Every tested pair was open-weight, so auditors who apply the method to closed models get the reader's view of the text and cannot measure that gap.

Publishers:lesswrong.com

Reality

Evidence35
Adoption
Insufficient
Hype gap+20
Incentives
Insufficient
Confidence30
build1 publisher

Deleting two terms from the tied-SAE gradient leaves the Oja update

A LessWrong post shows the Oja rule falling out of tied sparse-autoencoder descent once two terms are dropped, then reports a language-model test where an initialization change moved the backprop baseline more than the Hebbian gap it was meant to explain.

Publishers:lesswrong.com

Reality

Evidence46
Adoption
Insufficient
Hype gap+12
Incentives34
Confidence51