build1 publisher
One "be honest" line pulled a planted design flaw into GPT-5.6-Sol's abstract
A LessWrong write-up planted invalidating flaws in ML experiment logs and asked models to write the conference abstract. A second model scored the disclosure on three levels, and the published example is one before-and-after pair.
Publishers:lesswrong.com
Reality
- Evidence38
- Adoption
- Insufficient
- Hype gap+25
- Incentives35
- Confidence40