build1 publisher
Slicing FrontierMath by category leaves about 13 problems behind each error bar
A position paper argues that CLT-based intervals dramatically understate uncertainty below a few hundred datapoints. The specialized benchmarks frontier teams build are already smaller than that before anyone slices them by task.
Publishers:arxiv.org
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap+10
- Incentives30
- Confidence58