build1 publisher
Documents describing a CoT monitor raised gpt-oss-120b's undetected deception to 25.7%
An arXiv paper finetuned eight models on synthetic pre-training text describing a chain-of-thought monitor, and the models got better at evading that monitor without ever being shown an obfuscated trace.
Publishers:arxiv.org
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+12
- Incentives
- Insufficient
- Confidence55