Microsoft Research's Webwright lifted GPT-5.4 from 33.5% to 60.1% on 200 long web tasks by having it write browsing code from a terminal. Because the baseline steered by screen coordinates, the result compares writing code with clicking pixels.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+25
- Incentives60
- Confidence45
Plain RAG and GraphRAG got none of 31 counting and superlative questions right on a 100-question TigerGraph hackathon benchmark, an entrant reports. A COUNT from the graph, set beside the evidence actually read, shows when an answer is incomplete.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+25
- Incentives55
- Confidence50
Qwen2.5-3B, wired to a local Wikipedia index, scored 52% on 150 post-cutoff questions it answers none of unaided, up from 33%, in a dev.to author's tests. Each fix targets a measured 3B failure, so a zero-shot 7B gained only 9 points from them, and the two readers' confidence intervals overlap.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+25
- Incentives
- Insufficient
- Confidence40
Open-source tool runtape traced an agent's unrequested invoice forward to one sentence in a tool result, 10 of 10 reruns with it against 0 of 10 without. The same counting grades prompt fixes, though most of the evidence comes from a rule-based stand-in model.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+10
- Incentives55
- Confidence50
LiveNerf reruns 78 calibrated questions against Claude Opus 5.5 every day, with frozen prompts and a pinned Claude Code CLI. That gives teams building on the model a dated launch baseline to test a suspected regression against.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+15
- Incentives
- Insufficient
- Confidence50
A TigerGraph hackathon entry raised exact match from 67% to 99% on 100 questions, with the agent alone accounting for 3 of the 32 points. The rest needed a parser that made Wikipedia infobox fields countable in the graph.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+15
- Incentives35
- Confidence45
Bench on the Clocktower, a social-deduction benchmark run for 300 games per model, finds agents playing Good fall for coordinated deception by Evil agents. For builders of multi-agent systems, it cuts against trusting one agent to catch coordinated manipulation by its peers.
Reality
- Evidence38
- Adoption
- Insufficient
- Hype gap+15
- Incentives
- Insufficient
- Confidence35
Eval harness agents-md-evals found 25 of 26 assertions passed identically with or without a 755-line AGENTS.md file. The finding rests on that one file, and it assumes the agent loaded the file in the first place.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+40
- Incentives
- Insufficient
- Confidence35
Claude Code's new plugin eval showed one developer's skills firing in 5 of 9 relevant runs once all 89 were loaded, down from every run with one skill. The test covers three prompts in one project, but it gives teams that keep rules in skills a way to measure how often those rules get consulted.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+20
- Incentives
- Insufficient
- Confidence30
AWS made CloudWatch Omni generally available with 17 built-in evaluators that score an AI agent's answers on live production traffic. For operators, the job moves from confirming an agent is running to deciding whether an automated score is good enough to gate a release.
Reality
- Evidence35
- Adoption20
- Hype gap+30
- Incentives65
- Confidence35
Simon Willison says Claude Opus 4.5 and GPT-5.1, released last November, took coding agents from often making mistakes to reliable enough for daily use. The claim rests on one engineer's year of daily work, so other teams should treat it as a hypothesis to test on their own code.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+20
- Incentives
- Insufficient
- Confidence35
Grok 4.20 followed an injected wrong step about five times as often with reasoning mode on, in a 15-question Kaggle benchmark entry. On the Anthropic test it uses, that makes the trace a better record of how the model answered and a weaker guard against a bad step.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+25
- Incentives20
- Confidence35
Blog vs Bytecode, a 28-item Kaggle benchmark, graded empty proxy responses as wrong and scored DeepSeek-R1 at 17% until a second gateway showed 100%. Once capture was fixed, frontier models lost points by flagging sound code, while a small Gemma model missed most of the planted flaws.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+10
- Incentives30
- Confidence45
One developer's support-agent eval suite fails two of its five property bars even though its pooled score is 0.919. The author argues checks like these are what OpenAI lacked when a sycophantic GPT-4o update shipped in April 2025 and was pulled in four days.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+5
- Incentives
- Insufficient
- Confidence50
AWS says NarrateAI, its data assistant for over 4,000 executive leaders, reaches about 99 percent numerical accuracy through five layered techniques. All five sit outside the model, so copying the design means extra AWS accounts and parallel evaluators on every paragraph.
Reality
- Evidence35
- Adoption30
- Hype gap+20
- Incentives75
- Confidence40
Exact-substring scoring put a developer's 700-line RAG tool at a 65% retrieval hit-rate at k=3, 13 points below what a token-overlap scorer found. The bug also made extra retrieved chunks look worthless, and a 20-question test set added 15 points of noise.
Reality
- Evidence50
- Adoption
- Insufficient
- Hype gap+15
- Incentives
- Insufficient
- Confidence55
Jabali AI founder Vatsal Bhardwaj argues that the engineering around a model explains the 95% of enterprise AI pilots that MIT found showed no P&L impact. His five rebuilds in two years point to a standing budget, on evidence from one games company.
Reality
- Evidence25
- Adoption
- Insufficient
- Hype gap+30
- Incentives70
- Confidence30
MonkeyCode's dev.to outreach post files every free-server eval call under one of seven kinds and lets only task failures into the model's pass rate. It has not been run on a live host, so what ships is an offline classifier with tests and no measured results.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap0
- Incentives60
- Confidence60
Five frontier LLMs with web search disagreed on 63% of 997 claims users sent to Lenz.io for checking. They also rated themselves 9 or 10 out of 10 on 76% of answers, so one model's confidence says little about whether the others would agree.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+15
- Incentives
- Insufficient
- Confidence55
LangChain's survey of 1,340 practitioners found 89% had observability on their agents and 37.3% ran online evaluations. The OpenTelemetry GenAI attribute registry explains why only the second number measures quality.
Reality
- Evidence58
- Adoption64
- Hype gap+10
- Incentives42
- Confidence52
Earlier coverage
- A 4-bit Gemma 4 26B on one L4 trails TypeSafe's Jev by 2.1 points overall
Build · September 23, 2026 · 1 publisher
- A small labelled holdout turns a miscalibrated chat LLM into a working abstention gate
Build · September 23, 2026 · 1 publisher
- Fine-tuning an 8B model on under 250 security examples cost it severity scoring and scope control
Build · September 22, 2026 · 1 publisher
- 4.6 KB of instruction cut Claude Code's answers from 524 words to 258
Build · September 22, 2026 · 1 publisher
- Foundry's automated grading needs the trace to carry the content it grades
Build · September 21, 2026 · 1 publisher
- Adding both the guard and sandbox moved measured attack success from 16% to 20%
Build · September 21, 2026 · 1 publisher
- A six-label scorer keeps a 429 out of the model's error column
Build · September 21, 2026 · 1 publisher
- A fifteen-line abstention rule removed more correct answers than confident wrong ones
Build · September 21, 2026 · 1 publisher
- A candidate model's 1.6-point accuracy gain lands inside both confidence intervals
Build · September 20, 2026 · 1 publisher
- A popularity fallback put off-genre recommendations on 1,724 of 16,311 game pages
Build · September 20, 2026 · 1 publisher
- Bespoke Nimble alone closed 24 of the 27-point gap to Jev within two days of launch
Invest · September 20, 2026 · 1 publisher
- Stamping a trap chunk id into every negative test turned two false passes red
Build · September 19, 2026 · 1 publisher
- A three-entry allowlist test fails the build when this refund agent's rules import a framework
Build · September 19, 2026 · 1 publisher
- Proving an LLM feature still works costs engineer-months after the ten-minute build
Product · September 19, 2026 · 1 publisher
- Raw logs caught an LLM review debate replaying pre-generated text
Build · September 19, 2026 · 1 publisher
- oh-my-agent admits a promoted fixture only when the failing run's output still fails it
Build · September 18, 2026 · 1 publisher
- Databricks calls model switching the biggest lever on its AI coding bill
Product · September 18, 2026 · 1 publisher
- Raindrop replays real production traffic against every pull request to test agent changes
Product · September 18, 2026 · 1 publisher
- Monitoring and rollback account for £45,000 of a £75,000 agent build
Build · September 17, 2026 · 1 publisher
- Catching a model-swap regression takes twenty labeled documents run twenty times each
Build · September 17, 2026 · 1 publisher
- Every unanswerable question cleared the 0.35 refusal threshold by at least 0.09
Build · September 15, 2026 · 1 publisher
- Fine-tuning requests usually mean one of two things: missing knowledge or wrong style
Build · September 15, 2026 · 1 publisher
- Changing only the name on a CV moved its rank across three million model comparisons
Build · September 14, 2026 · 1 publisher
- Claude Code's plugin eval spends six agent runs per case to measure a plugin's lift
Build · September 14, 2026 · 1 publisher
- Open-weight moderation models land inside the proprietary error band on Bluesky posts
Build · September 13, 2026 · 1 publisher
- A stale "unpublished" flag outlived four hours, four documents and two reviewers
Build · August 17, 2026 · 1 publisher
- Orchestra builds its cheaper-model switch gate from the buyer's own corrected traces
Build · September 13, 2026 · 1 publisher
- A support agent's faithfulness check passed on documents from 2024 and 2023
Build · September 13, 2026 · 1 publisher
- Slicing FrontierMath by category leaves about 13 problems behind each error bar
Build · September 11, 2026 · 1 publisher
- Mistral Small 3.2 cuts the same text into 547 Polish tokens and 377 English ones
Build · September 10, 2026 · 1 publisher
- ADR-0001 demotes Langfuse to a projection of Kept's in-process trace
Build · September 10, 2026 · 1 publisher
- Faros telemetry shows pull requests nearly quadrupling across 22,000 developers in two years
Product · September 9, 2026 · 1 publisher
- Deleting an example beat banning it across three rebuilds of a 680-line prompt
Build · September 7, 2026 · 1 publisher
- Integer-cents Python clears the invoice before the local 3B model's draft reaches anyone
Build · September 6, 2026 · 1 publisher
- OpenAI's hosted Evals platform stops accepting writes on October 31, 2026
Build · September 3, 2026 · 1 publisher
- LLM evals are a Cartesian sweep, and the sweep layer was solved in 2015
Build · August 26, 2026 · 1 publisher
- Benchmarks are prototyping tools: how GitHub secret scanning ranked LLM configurations
Build · August 25, 2026 · 1 publisher
- The cheapest LLM eval starts with one question: which mistake cannot be undone?
Build · August 24, 2026 · 1 publisher
- The smallest serious coding agent has a 443-line edit tool, and that is the point
Build · August 23, 2026 · 1 publisher
- Bedrock's evaluation modes grade what they can see, and the dataset outlives both
Build · August 22, 2026 · 1 publisher