build1 publisher
CoT monitoring helps because reward shapes reasoning only indirectly, not because it works perfectly
A position paper argues that reasoning traces carry safety signal precisely because reinforcement learning treats them as latents rather than outputs, and it asks frontier developers to weigh training decisions against that.
Publishers:arxiv.org
Reality
- Evidence56
- Adoption
- Insufficient
- Hype gap−12
- Incentives55