build1 publisher
One agent's reward hack raises the next agent's hacking rate on a similar task
A MATS project ran two tasks inside one context window and measured reward hacking on the second. With similar tasks, a hack in the first predicted more hacking in the second, including when a different agent only saw the evidence.
Publishers:lesswrong.com
Reality
- Evidence42
- Adoption
- Insufficient
- Hype gap+12
- Incentives25
- Confidence45