Claude Code's new plugin eval showed one developer's skills firing in 5 of 9 relevant runs once all 89 were loaded, down from every run with one skill. The test covers three prompts in one project, but it gives teams that keep rules in skills a way to measure how often those rules get consulted.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+20
- Incentives
- Insufficient
- Confidence30
Timescale says Claude wrote nearly all the code and ran all 41 experiments that lifted its LoCoMo memory score from 0.392 to 0.666 F1 in six days. By the team's own account, the calls that decided the result came from a person checking whether each metric measured what it claimed to.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+5
- Incentives45
- Confidence40
OneFindMe's developer took axe-core to zero findings, then a keyboard and VoiceOver pass turned up six kinds of failure the scanner never flagged. One site and one blind iPhone user make a strong case for a screen reader in the release test plan.
Reality
- Evidence45
- Adoption8
- Hype gap+5
- Incentives25
- Confidence55
Throughline's on-call agent returns a three-way coverage verdict with every memory recall, so a timed-out search cannot pass as "no prior incidents". Its author hit the same error-as-empty trap in CockroachDB's managed MCP server, where failures come back as HTTP 200.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap−10
- Incentives30
- Confidence40
Loquent's 14,000 vox-bench runs found Claude Haiku's latency tail at 2.85x its median while ElevenLabs tightened to 1.92x. Alert thresholds tuned to June's data would now be watching the wrong stage of the pipeline.
Reality
- Evidence40
- Adoption15
- Hype gap+15
- Incentives55
- Confidence40
A Kaggle-challenge benchmark called ART scores models on whether they still flag a function after the fix is applied. On eight synthetic pairs, the difference between price tiers showed up only on the patched half.
Reality
- Evidence47
- Adoption12
- Hype gap−5
- Incentives58
- Confidence44
A 29-session test on Claude Code v2.1.273 ran the same protected-directory rule two ways, as prose in CLAUDE.md and as a PreToolUse hook. Both held on a plain task. The comparison that separates them rests on four runs.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+10
- Incentives
- Insufficient
- Confidence55
The company benchmarked coding agents on real tasks against its own multi-million line codebase and found that per-token price predicted almost nothing about what a finished task cost. GLM 5.2 came in at $1.28.
Reality
- Evidence58
- Adoption38
- Hype gap+20
- Incentives70
- Confidence55
Symfony AI Mate deleted its MCP server in August and now exposes profiler data through console commands that SKILL.md files tell the agent to run. The author of the community bundle Mate replaced has deprecated his.
Reality
- Evidence55
- Adoption68
- Hype gap+8
- Incentives55
- Confidence58
A developer pinned a model into all 11 of his own subagent files, then counted 576 launches and found 63% were built-ins inheriting the session default. One implementation plus one review emptied his top model's limit.
Reality
- Evidence62
- Adoption32
- Hype gap+12
- Incentives25
- Confidence58
A dev.to write-up fixes retrieval at 200-word chunks and cosine top-5, then swaps three embedders and five generators across 100 questions and four topics to find out whether its own evaluation conclusions hold.
Reality
- Evidence52
- Adoption15
- Hype gap−12
- Incentives35
- Confidence45
The loop reads Bedrock invocation logs through Athena, prices them at published rates, and denies Claude Opus at 80 percent of an engineer's daily budget. It transfers only if every model call carries a federated user identity.
Reality
- Evidence58
- Adoption30
- Hype gap+25
- Incentives78
- Confidence54
Webhands refuses any recipe containing a write-marked click until the caller resends it with confirm:true. That check runs before Cloudflare provisions a browser, so no mutation ever gets caught half-executed.
Reality
- Evidence42
- Adoption
- Insufficient
- Hype gap+16
- Incentives55
- Confidence38
Nvidia's CEO wants a $500,000 engineer burning $250,000 in tokens. That is half of payroll on one line item, and the companies that adopted the metric first are already writing caps.
Reality
- Evidence38
- Adoption58
- Hype gap+42
- Incentives76
- Confidence41
A dev.to build log puts a model in the browser at run time, then takes it out. The bill is the stated reason; the coin-flip results are the one that ends it for monitoring.
Reality
- Evidence30
- Adoption12
- Hype gap+28
- Incentives32
- Confidence45
An edge-hosted product search rewrote its prompt from "translate this query" to "name the product a seller would list". Haiku beat Sonnet once every search became an inference call.
Reality
- Evidence34
- Adoption16
- Hype gap+18
- Incentives62
- Confidence38
A Claude Code Stop hook appended 2,340 session rows that all said the session cost nothing, because it read fields the payload never carried. Flat-fee billing left no invoice to argue with it.
Reality
- Evidence44
- Adoption12
- Hype gap+22
- Incentives62
- Confidence38
A ShipWithAI post puts loop verification in order: exit code, Stop hook, second model. The hook's documented bypass and the missing cost number are where it gets interesting.
Reality
- Evidence34
- Adoption16
- Hype gap+16
- Incentives68
- Confidence38
Branden Jenkins found the bill on his phone at dinner. His staff hit approval limits before they can spend; his own token wallet just quietly refilled itself.
Reality
- Evidence38
- Adoption58
- Hype gap+18
- Incentives62
- Confidence45
A small Bedrock model sits between the retriever and the answer call and strips text the query does not need. Whether that pays comes down to two numbers AWS cannot supply for you.
Reality
- Evidence38
- Adoption
- Insufficient
- Hype gap+32
- Incentives86
- Confidence54