science1 distinct publisher
NIST puts two different error bars on the same AI benchmark score
AI 800-3 defines two accuracies a benchmark can estimate, one for the fixed question set and one for the population it stands for, and shows that the common grand-mean method understates confidence for the first.
Publishers:nist.gov
Reality
- Evidence57
- Adoption
- Insufficient
- Hype gap−14
- Incentives32