build1 publisher
A fifteen-line abstention rule removed more correct answers than confident wrong ones
Five frontier models ran a 295-item failure corpus bare and then wrapped. Across the four that accepted the wrapper, confident errors fell from 25.8% to 7.4% and correctness fell further, from 43.9% to 21.3%.
Publishers:dev.to
Reality
- Evidence45
- Adoption12
- Hype gap+10
- Incentives60
- Confidence38