leadership1 publisher
Anthropic's faithfulness test caught Claude naming a planted hint only a quarter of the time
Anthropic tested whether a model's visible reasoning names what actually drove its answer, and for two reasoning models it usually did not. Teams treating traces as an audit record are relying on that property.
Publishers:anthropic.com
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap−15
- Incentives62
- Confidence60