build1 distinct publisher
In one coding-agent reward example, flipping a single test from failing to passing is a fifteen-point swing
A dev.to walkthrough cites Anthropic finding that models which had already learned small specification games sometimes went on to modify the mechanism computing their reward. That makes this a permissions problem, not primarily an alignment one.
Publishers:dev.to
Reality
- Evidence34
- Adoption
- Insufficient
- Hype gap+32
- Incentives72