One developer's 'could not tell' exit code stopped a false green on a hold, then exposed four new ways a check can mislead. Each looks like a clean result, so agent evals that adopt the third value need controls that force every outcome.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap−5
- Incentives
- Insufficient
- Confidence50
Five agents built to cheat a ledger-scored refund benchmark found three scorer holes, including a $50 cap that scored $120 in split refunds as $0.00. An outage, too-good results and a fact-check found the other three.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap0
- Incentives30
- Confidence50
Huntress ran seven Rails tasks past Claude Fable 5.1 and watched it copy hand-rolled code out of its own repository. Changing the harness, in three cheap steps, took API recall from 48 percent to 100.
Reality
- Evidence58
- Adoption25
- Hype gap+15
- Incentives40
- Confidence55
A dev.to post credits Anthropic with running the same model and the same prompt under two harnesses, 20 minutes and $9 for a broken result against six hours and $200 for a working one. The hourly spend barely moved.
Reality
- Evidence25
- Adoption
- Insufficient
- Hype gap+30
- Incentives45
- Confidence40
AWS has added three skill-focused evaluators to Strands Evals and Bedrock AgentCore Evaluations. Two of them call a model once per invoked skill, and the third is a deterministic check that exists only in Strands.
Reality
- Evidence52
- Adoption10
- Hype gap+15
- Incentives80
- Confidence55
NVIDIA's developer blog sets out the five-level rollup from step to benchmark and the two scores that read the same trace. Step-level says where the chain broke. End-to-end reads the environment and says whether the refund posted.
Reality
- Evidence55
- Adoption20
- Hype gap+12
- Incentives50
- Confidence60
LangChain clocked TypeSafe AI's Jev at 0.44 seconds and $0.00035 per call against three LLM judges on the same eval set. Whether that price transfers depends on how much structure your traces already have.
Publishers:langchain.com
Reality
- Evidence45
- Adoption15
- Hype gap+18
- Incentives70
- Confidence55
A dev.to FAQ proposes failing any agent session whose diff touches a test path without the ticket ID in a signed allowlist. Both scripts in the post are labeled unexecuted, so the false-positive count is yours to find.
Reality
- Evidence34
- Adoption
- Insufficient
- Hype gap+14
- Incentives20
- Confidence50
COGEXT's author pushed 150 samples from cookbooks, DEV posts and Hacker News through his own extractor. The unbiased 120 yielded one promise, and the 13 in the enriched set averaged 0.79 confidence with a single deadline between them.
Reality
- Evidence38
- Adoption10
- Hype gap+24
- Incentives85
- Confidence58
License Referee resolves each npm dependency to an SPDX id and looks up a ruling for the exact dependency-into-project direction. On its 22-question test set, all three configurations agreed on the verdict and one invented its source.
Reality
- Evidence42
- Adoption12
- Hype gap+12
- Incentives72
- Confidence40
ProgramDistill turns 26 working web applications into thousands of verifiable tasks where the specification is the running software itself. Difficulty comes from stacking repairs that depend on one another, and the reported scores track it.
Reality
- Evidence52
- Adoption10
- Hype gap+18
- Incentives45
- Confidence50
A dev.to post lays out a coding-agent eval protocol that hashes the token and tool-call envelope alongside the tasks, so editing a cap changes the run id. No executed runs are published with it.
Reality
- Evidence34
- Adoption
- Insufficient
- Hype gap+12
- Incentives58
- Confidence56
Turing Post talked through agent feedback with Salesforce's chief AI scientist at Dreamforce. A correction lands in one of four places: the context, persistent memory, the software around the model, or the weights.
Publishers:turingpost.com
Reality
- Evidence46
- Adoption20
- Hype gap+12
- Incentives62
- Confidence48
A developer ran three change tasks in one Python and TypeScript product against two code-index tools, and counted a run only after the focused test passed and then failed again with the defect put back on purpose.
Reality
- Evidence46
- Adoption12
- Hype gap−15
- Incentives32
- Confidence41
Coding-agent pass rates get re-quoted after every prompt edit. A dev.to protocol freezes a seeded holdout split before any tuning, and at its minimum size the published number gets too coarse to show a 13-point gain.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+15
- Incentives20
- Confidence62
AWS's system prompt optimizer hands a reflector agent a shell tool and a directory of scored production traces, and offline batch evaluation plus a live-traffic A/B test decide which of its proposed edits reach the agent.
Reality
- Evidence45
- Adoption20
- Hype gap+15
- Incentives82
- Confidence45
A Microsoft developer blog post documents a Dev Proxy knowledge evaluation that blocked web tools and curl, then passed questions about recent versions. The agent had been reading a local source checkout.
Publishers:devblogs.microsoft.com
Reality
- Evidence55
- Adoption12
- Hype gap+12
- Incentives30
- Confidence45
The harness swaps a large tool response for a file path and a ten-line preview, truncates stale write arguments at 85 percent of the window, and summarises only when there is nothing left to move to disk.
Publishers:blog.langchain.com
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+10
- Incentives78
- Confidence62
Toolmetry edited only the strings that tell an agent what each MCP tool does and how to call it. On a single published run, its three test servers finished between 96.4 and 100 percent for $4 of API spend.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+55
- Incentives55
- Confidence38
The durable output of OpenAI's new agent cookbook is an eval suite generated from human and model feedback on five runs of one fictional company, plus a handoff file that tells Codex what to change next.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+20
- Incentives78
- Confidence62
Earlier coverage
- LangChain drops about 4,000 base input tokens from every default Deep Agents turn
Build · September 15, 2026 · 1 publisher
- A production agent's LLM judge filed 113 of 114 state defects under brand voice
Build · September 13, 2026 · 1 publisher
- Procedural Graphs keep an agent's flowchart edit only after it passes a validation set
Build · September 10, 2026 · 1 publisher
- Werewolf agents read the harness's own turn order as evidence of guilt
Build · September 10, 2026 · 1 publisher
- AWS's turn-level metric separates the one broken turn from the three that inherited it
Build · September 10, 2026 · 1 publisher
- Two canaries with known outcomes tell you whether an agent eval is scoring the model or the harness
Build · September 9, 2026 · 1 publisher
- Anthropic's multi-agent writeup puts the engineering weight on coordination and evaluation
Leadership · August 31, 2026 · 1 publisher
- AWS puts agent evaluation on OpenTelemetry, and on exactly three span roles
Build · August 26, 2026 · 2 publishers
- Anthropic cut 80% of Claude Code's system prompt and the evals did not move
Invest · August 23, 2026 · 1 publisher
- Harness-Bench makes the scaffold a measured variable, and model-only scores lose meaning
Build · August 23, 2026 · 1 publisher
- The missing field is the product: why agent listings need declared, derived and unknown
Build · August 23, 2026 · 1 publisher
- A SKILL.md layer quietly rerouted an agent off the MCP tools it was given
Build · August 22, 2026 · 1 publisher
- Your agent didn't misunderstand the prompt. It ran the wrong branch.
Build · August 16, 2026 · 1 publisher