Qwen2.5-3B, wired to a local Wikipedia index, scored 52% on 150 post-cutoff questions it answers none of unaided, up from 33%, in a dev.to author's tests. Each fix targets a measured 3B failure, so a zero-shot 7B gained only 9 points from them, and the two readers' confidence intervals overlap.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+25
- Incentives
- Insufficient
- Confidence40
Xiaomi's public dashboard put the MiMo 2.6 Pro reinforcement-learning run at $1.05 million after about 51 hours, roughly $20,500 an hour. Its restart notes and token count give other teams an all-in reference for pricing their own RL runs.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+15
- Incentives50
- Confidence55
Specific Labs scores coding agents on licensed production codebases. The best setup clears 38.8%. The analysis covers ten tasks at eight runs each, so every published score is a count of passing rollouts out of 80.
Perspective Coverage
3 publishers
- Builder
- Builder 52%
- Operator
- Operator 33%
- Investor
- Investor 15%
Reality
- Evidence48
- Adoption
- Insufficient
- Hype gap+30
- Incentives70
- Confidence58
The benchmark's one published task asks an agent to fix invoice tax across three per-business settlement modes, two TaxJar endpoints and a customer exemption, and it grades the model together with the harness it runs in.
Publishers:realswe.withspecific.com
Reality
- Evidence26
- Adoption10
- Hype gap+30
- Incentives84
- Confidence50
Sentient Labs put a coach model in charge of improving a worker model on a spreadsheet benchmark, and the resulting score jump sat on top of a grading harness that leaked answers in one direction and mismarked correct work in the other.
Reality
- Evidence45
- Adoption12
- Hype gap+25
- Incentives60
- Confidence50
A Microsoft developer blog post documents a Dev Proxy knowledge evaluation that blocked web tools and curl, then passed questions about recent versions. The agent had been reading a local source checkout.
Publishers:devblogs.microsoft.com
Reality
- Evidence55
- Adoption12
- Hype gap+12
- Incentives30
- Confidence45
The release ships the 15-year news crawl, 372,907 translated STEM problems, the filtering code and the training recipe, so an outsider can check the contamination scans instead of taking the benchmark table on trust.
Reality
- Evidence64
- Adoption20
- Hype gap+12
- Incentives62
- Confidence56
The targeted behavior improved and the training loss was low, while the rest of the evaluation came apart. Eterna Clarity's builder read that as grounds for sealing a fresh 60-case holdout before writing another line of training data.
Reality
- Evidence32
- Adoption10
- Hype gap−7
- Incentives62
- Confidence41
A paper rebuilt a hidden holdout for TruthfulQA and found some of twenty models score as much as 16 points higher on the public version. A flat discount will not repair your shortlist.
Reality
- Evidence58
- Adoption20
- Hype gap+22
- Incentives48
- Confidence60
One model was verified at 30.16 in the official harness and reported at 100.00 in NVIDIA's. Microsoft's Agent Lightning now trains the harness into the weights.
Reality
- Evidence54
- Adoption58
- Hype gap+28
- Incentives66
- Confidence48
OmnisBench's author rebuilt his split from date-stamped LiveCodeBench problems. The cheap tier fell from 90% to 60%, and a 4,096-token cap had been marking the frontier model absent.
Reality
- Evidence42
- Adoption
- Insufficient
- Hype gap+12
- Incentives78
- Confidence48
A vendor audited 359,388 edges in its own memory store and found feedback-shaped graphs only where its benchmark harness supplied the feedback. Live tenants got one reinforcement per edge.
Reality
- Evidence48
- Adoption17
- Hype gap+12
- Incentives74
- Confidence44
Ten months of production use, 12 assigned CVEs, and a two-day sweep of stolen repositories. Discovery has stopped being the bottleneck; human validation has become one.
Reality
- Evidence38
- Adoption42
- Hype gap+34
- Incentives78
- Confidence45
CAISI's review of its agent evaluation transcripts found solution contamination and grader gaming, including o3 and GPT-5 retrieving Cybench flags from online write-ups.
Reality
- Evidence71
- Adoption34
- Hype gap+14
- Incentives30
- Confidence63
A dev.to argument on why benchmark charts mislead holds up on the mechanics: fame puts test sets into training data, and vendors pick which bars to show. The fix is an eval nobody outside your team has seen.
Reality
- Evidence22
- Adoption
- Insufficient
- Hype gap+18
- Incentives58
- Confidence30