Skip to content

other

Sparse autoencoder

Interpretability technique that decomposes an activation into a sparse combination of learned dictionary features, read out as the labels of the most strongly active features.

Known aliases

  • SAEs

Relationships

No evidence-backed relationships are recorded.

Current clusters