build1 publisher
Repeat runs rule out chance as the source of position bias in twelve LLM judges
A study on arxiv pushes more than 100,000 evaluation instances through 12 LLM judges on MTBench and DevBench, and reports that the quality gap between two candidate answers moves position bias far more than prompt length does.
Publishers:arxiv.org
Reality
- Evidence68
- Adoption
- Insufficient
- Hype gap+10
- Incentives30
- Confidence58