Simon Willison says Claude Opus 4.5 and GPT-5.1, released last November, took coding agents from often making mistakes to reliable enough for daily use. The claim rests on one engineer's year of daily work, so other teams should treat it as a hypothesis to test on their own code.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+20
- Incentives
- Insufficient
- Confidence35
DigitalOcean says step-by-step reasoning is billed as output and invisible by design. The 90% figure it cites comes from a paper that estimates hidden token counts, so moving it onto your own invoice takes a matching task mix.
Reality
- Evidence32
- Adoption26
- Hype gap+38
- Incentives88
- Confidence42
Anthropic's open-source audit framework now runs a classifier over every auditor turn and rewrites anything a real deployment would not produce. The tuning targeted models that say out loud they are being tested.
Publishers:alignment.anthropic.com
Reality
- Evidence55
- Adoption35
- Hype gap−10
- Incentives75
- Confidence45
Microsoft Research says a 4-billion-parameter model trained with SocialRL out-scored GPT-4.1, GPT-5.1 and GPT-5.2 across six bargaining domains. The winning margin is thinner than the method's effect.
Reality
- Evidence28
- Adoption
- Insufficient
- Hype gap+46
- Incentives68
- Confidence33
A Secure Code Warrior and RMIT study of six frontier models across 11 frameworks found no universal winner and no link between token cost and secure output.
Reality
- Evidence48
- Adoption
- Insufficient
- Hype gap+20
- Incentives68
- Confidence40