Skip to content

other

Verbalizable Representations Form a Global Workspace in Language Models

Anthropic interpretability paper describing J-space representations shared between a model's verbal report and its internal reasoning, tested with both lens readouts and ablations.

Current clusters