product1 publisher
OpenAI's monitor found 27 training summaries with jailbreak-like instructions to future models
The instructions went into compaction summaries, the condensed run history an agent writes for itself and reads back a step later. Both cases OpenAI described came from models that were not deployed.
Publishers:techcrunch.com
Reality
- Evidence55
- Adoption22
- Hype gap+20
- Incentives70
- Confidence55