Skip to content

Topic

Sparse autoencoders

Models trained to reconstruct neural network activations through an overcomplete dictionary where only a few latents fire at once, used to decompose representations into individually interpretable directions.

Current clusters

build1 publisher

Deleting two terms from the tied-SAE gradient leaves the Oja update

A LessWrong post shows the Oja rule falling out of tied sparse-autoencoder descent once two terms are dropped, then reports a language-model test where an initialization change moved the backprop baseline more than the Hebbian gap it was meant to explain.

Publishers:lesswrong.com

Reality

Evidence46
Adoption
Insufficient
Hype gap+12
Incentives34
Confidence51