security1 publisher
Petri 2.0 screens its own auditor to keep models from noticing they are under test
Anthropic's open-source audit framework now runs a classifier over every auditor turn and rewrites anything a real deployment would not produce. The tuning targeted models that say out loud they are being tested.
Publishers:alignment.anthropic.com
Reality
- Evidence55
- Adoption35
- Hype gap−10
- Incentives75
- Confidence45