build1 publisher
Polling-style aggregation often amplified shared misconceptions across five benchmarks
An arXiv preprint ran majority voting, confidence weighting and the Surprisingly Popular algorithm over several open models, and the errors were correlated enough that none of them beat a single sample.
Publishers:arxiv.org
Reality
- Evidence52
- Adoption
- Insufficient
- Hype gap+12
- Incentives25
- Confidence45