product1 publisher
OpenAI's Daniel Selsam says models are getting too situationally aware to evaluate
Selsam published his warning the same week two alignment researchers left Anthropic and Google DeepMind. His specific claim is that evaluation results are losing their power to predict how a model behaves when it is not being watched.
Publishers:gizmodo.com
Reality
- Evidence42
- Adoption
- Insufficient
- Hype gap+38
- Incentives62
- Confidence48