MIT and Sakana AI's SIFT ran a coding agent's full self-improvement search on roughly $34 of API calls, about a tenth of the Darwin Godel Machine's resources. The saving comes from a language-model judge screening patches, so it holds only while that judge picks correctly.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+25
- Incentives
- Insufficient
- Confidence45
Apollon Labs' Greek benchmark found gpt-oss-20b making up 48 words per 1,000, against zero for six of the 15 models tested. It targets a hallucination that fact checks skip: words that look Greek and do not exist.
Reality
- Evidence48
- Adoption
- Insufficient
- Hype gap+5
- Incentives35
- Confidence50
Retirement-answer-check's injection regex caught all 16 first-round attacks and none of the 20 written by a second red team that had read it. Its two model judges, told to treat drafts as untrusted data and given an injection flag, caught all 16 of the new attacks in every run.
Reality
- Evidence55
- Adoption3
- Hype gap+15
- Incentives30
- Confidence50
CDV's adaptive stop rule truly reached its quality bar in 71% of runs under judge noise of 0.10, against 93% for a fixed six-step budget. Requiring two passing scores in a row lifts that to 97% for 1.4 extra steps, so a one-line change beats both the Bayesian policy and the budget.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+8
- Incentives30
- Confidence55
A PyTorch case study credits Shopify's GraphQL agent with beating frontier-model quality at 96 percent lower serving cost. Most of what it documents is how the reward signal behind that gets defined and checked.
Reality
- Evidence45
- Adoption35
- Hype gap+30
- Incentives75
- Confidence50
A Magic: The Gathering rules agent works out which source is entitled to settle a question before it answers. Its evaluation pits 496 structured documents against a single BM25 pass over the same corpus, graded blind.
Reality
- Evidence45
- Adoption12
- Hype gap−10
- Incentives65
- Confidence50
With documentation and web search removed, GPT-5.6 Luna passed 18% of 336 Dev Proxy tasks and 15% of 413 SPFx tasks, and the passes appear throughout both product histories instead of stopping at one release.
Publishers:devblogs.microsoft.com
Reality
- Evidence62
- Adoption18
- Hype gap+12
- Incentives55
- Confidence55
A dev.to developer prices each API request after the work by asking a second model how much the answer represents, and the first version threw away the calibration he was paying for by rounding its probabilities into three tiers.
Reality
- Evidence38
- Adoption8
- Hype gap−12
- Incentives45
- Confidence52
CommentBench splits human comments on AI-safety posts and drafts into target points, filters out the ones a model could not reach without extra context, and has Opus 5 judge which of the rest a model hit.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+15
- Incentives55
- Confidence50
The four RAGAS-lineage metrics were built to separate a retriever's failures from a generator's. Read in pairs, they also expose the case where the model skipped the context and got the answer right anyway.
Reality
- Evidence22
- Adoption
- Insufficient
- Hype gap+35
- Incentives
- Insufficient
- Confidence35
A LessWrong write-up planted invalidating flaws in ML experiment logs and asked models to write the conference abstract. A second model scored the disclosure on three levels, and the published example is one before-and-after pair.
Reality
- Evidence38
- Adoption
- Insufficient
- Hype gap+25
- Incentives35
- Confidence40
Each case in claude plugin eval runs three times with the plugin loaded and three times without it. That doubling is what produces the delta column, and every one of those calls is billed to your own credentials.
Publishers:code.claude.com
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap0
- Incentives72
- Confidence66
Cloudflare says its graph model flagged eight malicious scripts running on live storefronts. A retrospective check put seven of the eight outside VirusTotal entirely. URLScan flagged none of the eight as malicious.
Reality
- Evidence45
- Adoption40
- Hype gap+30
- Incentives85
- Confidence50
A dev.to post publishes an unexecuted sketch that splits deterministic structural checks from a model-graded rubric and pins both graders to a changelog file, so a case citing a missing version stops the run before anyone reads a pass rate.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+12
- Incentives18
- Confidence55
Its engineering post argues that automated tests for multi-turn tool use belong in place before an agent scales. The evidence behind the advice is three deployments, one of them Anthropic's own.
Reality
- Evidence45
- Adoption35
- Hype gap+18
- Incentives78
- Confidence55
The 26% clean-fix rate now quoted to keep coding agents out of patch work came from six hand-picked hard bugs, and 22% of the runs behind it instructed the agent to apply a fix known to be wrong.
Reality
- Evidence62
- Adoption20
- Hype gap+12
- Incentives68
- Confidence55
A developer who built his own eval gate reran an unchanged 21-case suite three times in ten minutes and got three different scores, with the run files recording identical prompt checksums each time.
Reality
- Evidence60
- Adoption20
- Hype gap−15
- Incentives40
- Confidence55
A three-model panel plus a judge plus a synthesis pass costs roughly four to five times one completion and often runs two to three times slower. With the bare model slug, the timing of that spend goes unrecorded in your repo.
Reality
- Evidence50
- Adoption
- Insufficient
- Hype gap+20
- Incentives70
- Confidence55
Exhaustive human review of a deployed food ordering agent found that its judge writes the defect down and then scores it on the wrong axis. The shipping gate trips only on hangs and hard assertions, so 100 rounds passed clean.
Reality
- Evidence58
- Adoption20
- Hype gap+14
- Incentives40
- Confidence55
The headline confirm rate scored agreement with the scanner's claim and bug detection in one number, so the follow-up reruns the same 200 OWASP slices with the flag removed and the predictions committed first.
Reality
- Evidence48
- Adoption15
- Hype gap+8
- Incentives35
- Confidence45
Earlier coverage
- AWS's turn-level metric separates the one broken turn from the three that inherited it
Build · September 10, 2026 · 1 publisher
- Anthropic broke an agent ceiling by making "is this design good?" a gradable question
Leadership · September 10, 2026 · 1 publisher
- Judge model choice swings AI-Infra-Guard's false positive rate fifteenfold
Security · September 9, 2026 · 1 publisher
- Six calls, eleven evaluators, and the pass condition that still lets dead air through
Build · August 26, 2026 · 1 publisher
- Three agents, one spec, and a blind judging round that mostly proved they vote for themselves
Build · August 22, 2026 · 1 publisher
- Palomar registers Lean proofs against a commit, and refuses to referee them
Build · August 18, 2026 · 1 publisher