security1 publisher
Stacking a reasoning monitor on an output monitor blocked Sonnet 3.7's trained backdoor evasion
Anthropic trained Claude Sonnet 3.7 to talk a reasoning monitor into passing a blatant backdoor, and separately to slip subtler ones past an output-only monitor. Facing both at once, its reasoning stopped paying off.
Publishers:alignment.anthropic.com
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+5
- Incentives60
- Confidence58