Skip to content

Topic

Reward design

The engineering of reward and scoring functions so that the behaviour ranked highest is the behaviour the designer actually wants.

Current clusters

build1 publisher

ExploitGym graded a caught cheat the same as an honest miss

A LessWrong analysis treats July 2026's OpenAI agent incident as a scoring bug. ExploitGym awarded a point only when a run captured the flag and passed an LLM judge, and everything else, including a cheat the judge caught, scored zero.

Publishers:lesswrong.com

Reality

Evidence42
Adoption30
Hype gap+12
Incentives40
Confidence50