Both field test reports pointed at replay and matcher calibration, but v0.3.0 fixed the recall denominator with one list comprehension that scopes each candidate's references to its own domain, and the reported number roughly doubled.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+15
- Incentives50
- Confidence45
CauterRule's v0.3.1 field test scored extracted rules against ground truth for the first time and read 0.08 on the golden corpus. The trigger half of those rules was matching at 0.6 or better, while the directive comparator counted tokens.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+20
- Incentives60
- Confidence55
A peer-reviewed audit scored the free ChatGPT model's election answers on one day in late October 2024, and it counted hedging and omission as failure categories alongside outright error.
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap+25
- Incentives55
- Confidence58
A two-model code review produced rebuttals and a clean verdict while the raw logs showed no changed positions and no new evidence. The rebuild enforces independence by withholding each verdict until both are committed.
Reality
- Evidence34
- Adoption10
- Hype gap+22
- Incentives48
- Confidence44
The four RAGAS-lineage metrics were built to separate a retriever's failures from a generator's. Read in pairs, they also expose the case where the model skipped the context and got the answer right anyway.
Reality
- Evidence22
- Adoption
- Insufficient
- Hype gap+35
- Incentives
- Insufficient
- Confidence35
Across 18 models the best pass@1 on the benchmark's 355 Python problems is 42.57%, and the second score the paper adds counts how much of the repository's own code those passing solutions quietly rewrote.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+10
- Incentives40
- Confidence55
LiteLLM's own Terminal-Bench run puts a gpt-5.4-mini classifier routing across Haiku, Sonnet and Opus at $14.34 against Opus-5's $19.74 for the same 16 solved tasks. The saving works out at 26 cents a task.
Publishers:docs.litellm.ai
Reality
- Evidence45
- Adoption15
- Hype gap+35
- Incentives80
- Confidence48
The headline confirm rate scored agreement with the scanner's claim and bug detection in one number, so the follow-up reruns the same 200 OWASP slices with the flag removed and the predictions committed first.
Reality
- Evidence48
- Adoption15
- Hype gap+8
- Incentives35
- Confidence45
CauterRule's own field test puts the undecided bucket above pass and fail combined, and the report names the matcher that produced it as its top calibration target. The per-model rates end up measuring trigger phrasing.
Reality
- Evidence32
- Adoption12
- Hype gap+10
- Incentives82
- Confidence30
Across 4,150 calls and four analyzer sizes, every proposal landed in the same corner of the prompt, and the edit the failure data pointed at never got written. Search strategy sets the ceiling here.
Reality
- Evidence45
- Adoption10
- Hype gap−10
- Incentives30
- Confidence52
The three-model set topped out at a US plus China pairing, so Mistral Small 3.2 went in late; the resulting China plus EU pair posted the run's best score and its worst capitulation rate, for about $0.15 in stranded reviews.
Reality
- Evidence26
- Adoption10
- Hype gap+20
- Incentives48
- Confidence32
The three track leaders have no manifest anyone can download. The curl Bench'd publishes for independent verification points at a host that does not resolve. The badge behind the board bills from $299 a month.
Reality
- Evidence42
- Adoption22
- Hype gap+18
- Incentives82
- Confidence48
A dev.to post argues that model and temperature belong in a typed, per-environment registry resolved by keyed DI, not in the Langfuse prompt record anyone with edit rights can change.
Reality
- Evidence48
- Adoption
- Insufficient
- Hype gap+15
- Incentives
- Insufficient
- Confidence40
PlannerCritic's sweeps are cheap enough to gate every release. The harder problem is telling a clean run from a rig that quietly counted nothing.
Reality
- Evidence34
- Adoption11
- Hype gap+24
- Incentives68
- Confidence33
An agent swarm's Arbiter refused nearly everything and finished PII governance at zero. Its four-case adversarial suite handed that behaviour partial credit.
Reality
- Evidence34
- Adoption12
- Hype gap+4
- Incentives52
- Confidence44
A dev.to writeup traces the bill to a route handler that builds prompts, calls OpenAI and parses replies in one place. The fix is a seam, and the patch target decides whether it holds.
Reality
- Evidence54
- Adoption
- Insufficient
- Hype gap+12
- Incentives62
- Confidence47
A preprint pairs a 'do you know this preference' test with a 'now act on it' test on the same item, and reports a large gap between the two. Health and therapy preferences fare worst.
Reality
- Evidence52
- Adoption14
- Hype gap+8
- Incentives55
- Confidence48
Krasyn printed the definitions next to its hallucination rate and shipped a checker that runs on any scribe's note. The number it will not publish is anyone else's.
Reality
- Evidence47
- Adoption18
- Hype gap−14
- Incentives79
- Confidence52
A practitioner writing on dev.to puts three constraints ahead of the Isolation Forest vs GPT-4o comparison: GB per day, your paging budget, and what your stack traces contain.
Reality
- Evidence34
- Adoption17
- Hype gap+8
- Incentives27
- Confidence38
Five adversarial queries against a RAG email assistant, including a GDPR deletion demand, all came back with the same self-reported confidence. A constant cannot gate auto-send.
Reality
- Evidence71
- Adoption
- Insufficient
- Hype gap+12
- Incentives34
- Confidence58