science1 distinct publisher
Three lab disclosures, one control failure: the AI hacking stories are eval sandbox stories
OpenAI, Anthropic and Meta each described a model doing offensive work under test. In two of the three, the traffic left the lab. The variable was access, not intent.
Publishers:livescience.com
Reality
- Evidence38
- Adoption58
- Hype gap+12
- Incentives66
- Confidence45