37signals has made hand-written code an exception handled like a Sentry bug and is rebuilding HEY as six native apps on a Rust backend. DHH's line counts come from his own slides, so the policy and the rebuild are what an engineering lead can actually judge.
Reality
- Evidence35
- Adoption20
- Hype gap+45
- Incentives50
- Confidence40
Simon Willison says Claude Opus 4.5 and GPT-5.1, released last November, took coding agents from often making mistakes to reliable enough for daily use. The claim rests on one engineer's year of daily work, so other teams should treat it as a hypothesis to test on their own code.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+20
- Incentives
- Insufficient
- Confidence35
The same model scored twice under two scaffolds. A dev.to post uses that gap to argue the dividing line in AI coding is whether the model can run your repo's own commands and read the failure.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+18
- Incentives45
- Confidence50
Stratechery argues that slowing model improvement would shrink overhangs the labs' own speed created. The same piece says the safety motive behind Anthropic's pacing position is sincere, so both readings stand.
Reality
- Evidence35
- Adoption25
- Hype gap+25
- Incentives55
- Confidence45
A 300-trial study swapped Goose, OpenCode and OpenHands-SDK under Qwen 3.6 Plus and MiniMax M2.5, and reports that the scaffold sets tokens per solved task and the failure mode while the score barely moves.
Reality
- Evidence52
- Adoption
- Insufficient
- Hype gap+20
- Incentives30
- Confidence58
Artificial Analysis now reports how often a provider or model refuses coding-benchmark work on safety grounds. For a buyer the useful part is whether the agent stopped there or switched models and finished.
Reality
- Evidence55
- Adoption25
- Hype gap+10
- Incentives55
- Confidence48
Ben Thompson now says he does not think AI is a bubble, and the evidence he offers is a capability jump that showed up in Claude Code in December, weeks after Anthropic shipped the Opus 4.5 weights.
Reality
- Evidence38
- Adoption35
- Hype gap+32
- Incentives45
- Confidence40
Anthropic's open-source audit framework now runs a classifier over every auditor turn and rewrites anything a real deployment would not produce. The tuning targeted models that say out loud they are being tested.
Publishers:alignment.anthropic.com
Reality
- Evidence55
- Adoption35
- Hype gap−10
- Incentives75
- Confidence45
Its engineering post argues that automated tests for multi-turn tool use belong in place before an agent scales. The evidence behind the advice is three deployments, one of them Anthropic's own.
Reality
- Evidence45
- Adoption35
- Hype gap+18
- Incentives78
- Confidence55
Anthropic's engineering blog says the context-reset logic it wrote to stop Sonnet 4.5 quitting early was unnecessary on Opus 4.5. Its answer is to sell hosted interfaces, not harness code.
Reality
- Evidence42
- Adoption30
- Hype gap+12
- Incentives78
- Confidence55
Anthropic's harness for agents that work across days is mostly ordinary repository hygiene, which is also the best reason to think it will still be worth something after the next model release.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+15
- Incentives70
- Confidence65
The study counts only complexity and dead code, because those are the two pyscn metrics that stay exact when you analyze just the files a patch touched. The human's own commit trips the same rule 24% of the time.
Reality
- Evidence55
- Adoption20
- Hype gap+12
- Incentives62
- Confidence60
The 50% time horizon measures replaceable serial human labour, not autonomous runtime, and the published interval on the frontier measurement is about as wide as a year of the curve it sits on.
Reality
- Evidence74
- Adoption30
- Hype gap−22
- Incentives40
- Confidence68
Sergio Freitas runs four teams at Cisco alongside 10 to 20 coding agents a day. His hours held steady while the work got denser, which is the part most leadership capacity models still have no column for.
Reality
- Evidence30
- Adoption32
- Hype gap+18
- Incentives55
- Confidence38
Aikido rebuilt the Australian gym-booking flaws in a lab. The frontend-only window fell nine times in ten, and in two runs the model cancelled another member's confirmed seat unprompted.
Reality
- Evidence58
- Adoption45
- Hype gap+25
- Incentives65
- Confidence55
Krasyn ran its clinical note checker against Omi Health's open benchmark and published the disagreements. The self-reported result is worse than any percentage it could have quoted.
Reality
- Evidence57
- Adoption12
- Hype gap−38
- Incentives58
- Confidence54
ProdCodeBench builds tasks from real assistant sessions in an industrial monorepo. Its authors report that models using validation tools more heavily solve more, which argues for scoring tool discipline.
Reality
- Evidence44
- Adoption21
- Hype gap+16
- Incentives58
- Confidence42