Skip to content

Topic

LLM interpretability

Research into how large language models represent information internally, including probing hidden states and identifying directions in activation space that correspond to particular behaviours.

Current clusters