Microsoft and Hugging Face's ThinkingBox found that 67.24% of 79,853 failed agent runs ended cleanly, with no final tool error. Those failures showed up only when executable checks read the records each run left in the backend.
Perspective Coverage
3 publishers
- Builder
- Builder 52%
- Operator
- Operator 38%
- Investor
- Investor 10%
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+10
- Incentives35
- Confidence68
Only a renamed tool argument got through Strands, LangGraph and CrewAI cleanly in a 36-run schema-change test posted on dev.to. Harsher changes let Strands and CrewAI exit 0 with nothing verified while LangGraph crashed outright, so each framework needs its own schema-change test.
Reality
- Evidence42
- Adoption
- Insufficient
- Hype gap+12
- Incentives15
- Confidence40
A dev.to post puts the fragile layer of a multi-tool agent in the choice between two tools that both fit the request, and argues that better tool descriptions plateau because the ambiguity belongs to the request itself.
Reality
- Evidence22
- Adoption
- Insufficient
- Hype gap+12
- Incentives22
- Confidence34
IBM Research's ALTK-Evolve distils an agent's own trajectories into scored guidelines and injects the top five at inference time, and a companion post puts a number on the reliability an average success rate hides.
Reality
- Evidence45
- Adoption20
- Hype gap+15
- Incentives85
- Confidence55
IBM Research calls the 24-point drop between passing once and passing five times the consistency gap. Independence would have predicted a 27% five-run rate, so the failures are clustering on particular tasks.
Reality
- Evidence45
- Adoption20
- Hype gap+12
- Incentives45
- Confidence48
Every task in the benchmark is checked against its required terminal state, so any tool order can pass and any extra side effect fails. Claude Opus 5, the strongest model tested, cleared 66.50% at pass@1.
Reality
- Evidence60
- Adoption18
- Hype gap+12
- Incentives35
- Confidence52
One team's canonical records said a piece was unshipped for 260 minutes after it went live. The rule about checking the artifact only runs in one direction, and that is the cheaper direction.
Reality
- Evidence45
- Adoption18
- Hype gap+18
- Incentives35
- Confidence55
A failed Fivetran lookup fell back to a cached snapshot, and deterministic logic promoted last-known state into a current Healthy verdict. The request returned successfully.
Reality
- Evidence44
- Adoption
- Insufficient
- Hype gap+12
- Incentives62
- Confidence41
A new benchmark penalises agents for acting on facts the user already revoked. The dedicated memory layers barely beat the plain models on it.
Reality
- Evidence42
- Adoption10
- Hype gap+22
- Incentives62
- Confidence45
A 58-day log of 78 unattended agents found 43% of failures were malformed output returning HTTP 200. On one day uptime read 97-100% while finished deliverables were zero.
Reality
- Evidence56
- Adoption24
- Hype gap+14
- Incentives44
- Confidence54
A dev.to writeup makes the structural case: tool and schema drift exists only as a diff between two snapshots, so no single health probe can find it, however thorough.
Reality
- Evidence52
- Adoption10
- Hype gap+16
- Incentives78
- Confidence44
Upstage is selling tool-calling discipline rather than reasoning, and says 370 billion tokens moved through OpenRouter in its first week. The price, the part that matters most, is still qualitative.
Reality
- Evidence27
- Adoption40
- Hype gap+33
- Incentives83
- Confidence41
A controlled evaluation across four benchmarks found centralized coordination lifted financial reasoning 80.9%, while every multi-agent variant tested made strict sequential planning worse.
Reality
- Evidence46
- Adoption18
- Hype gap+12
- Incentives62
- Confidence52
A dev.to write-up splits agent work into prompts, context and harness. The interesting part is the harness: tool execution, permissions, validation and recovery, all of it code you own.
Reality
- Evidence26
- Adoption
- Insufficient
- Hype gap+22
- Incentives34
- Confidence38