A reader's rule said to suspect the test set before the methods, and checking it showed the fine-tune was the arm the easy data flattered most, losing 33 points on the rebuilt set against prompting's 28.
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap−12
- Incentives34
- Confidence58
A three-arm rerun of a word-length cipher puts positional agreement at 0.175 against a 0.044 random floor, with a prior-only control at 0.129 whose interval touches the treatment. That gap is where most self-consistency numbers live.
Reality
- Evidence46
- Adoption
- Insufficient
- Hype gap+17
- Incentives22
- Confidence56
A paper rebuilt a hidden holdout for TruthfulQA and found some of twenty models score as much as 16 points higher on the public version. A flat discount will not repair your shortlist.
Reality
- Evidence58
- Adoption20
- Hype gap+22
- Incentives48
- Confidence60
Cisco Talos reports that turning reasoning effort up often cost more without scoring better, and sometimes scored worse. The number that should drive procurement is the worst run, not the median.
Reality
- Evidence58
- Adoption15
- Hype gap+10
- Incentives55
- Confidence48
A dev.to writeup measured what an unchanged prompt does to itself over the same inputs, and the answer makes most six-item side-by-side comparisons unreportable.
Reality
- Evidence34
- Adoption
- Insufficient
- Hype gap+18
- Incentives22
- Confidence33
Stanford's April 2026 tests put debate ahead of rival multi-agent designs, and behind a solo agent at matched compute. The comparison a framework vendor quotes is the first one.
Reality
- Evidence30
- Adoption16
- Hype gap
- Insufficient
- Incentives
- Insufficient
- Confidence32
The acc and acc_norm split in lm-eval-harness can move in opposite directions on one checkpoint. Pick the metric before you train, and say which one you picked.
Reality
- Evidence52
- Adoption
- Insufficient
- Hype gap+14
- Incentives18
- Confidence46
Arize and Fireworks ran ten models against 40 agent tasks and found the cheapest model per finished job also had the worst pass rate. Coverage, not price, is the binding constraint.
Publishers:arize.com
Reality
- Evidence52
- Adoption20
- Hype gap+22
- Incentives78
- Confidence45
An arXiv preprint argues single agents are more information-efficient under a fixed reasoning budget, and reports they match or beat orchestrated agents on multi-hop reasoning.
Reality
- Evidence52
- Adoption
- Insufficient
- Hype gap+12
- Incentives
- Insufficient
- Confidence44
A verify-on-read experiment rerun across 14 live models on a fingerprinted 50-fact set found false-accept rates up to 0.38, and run-to-run noise wide enough to swallow a prompt fix.
Reality
- Evidence57
- Adoption14
- Hype gap−12
- Incentives31
- Confidence44