product1 publisher
OpenAI counts 27 training summaries in which a model wrote jailbreaks for its own future self
The lab published six misalignment incidents from the past six months plus a commitment to release future ones before it has explained or fixed them. Three of the six end with a model inventing data it could not fetch.
Publishers:gizmodo.com
Reality
- Evidence48
- Adoption25
- Hype gap+12
- Incentives74
- Confidence55