build1 publisher
DeepMind's interpretability lead found Anthropic's J-space again in an open 27B model
Anthropic screened Claude for one property, verbalizability, and found two more in the same representations. Neel Nanda reproduced the structure in open weights, and the outside commentators Anthropic invited disagree about what it is.
Publishers:lesswrong.com
Reality
- Evidence40
- Adoption22
- Hype gap+12
- Incentives62
- Confidence42