Opus 5.5 lists at twice Sonnet 5's price, yet in one developer's matched Claude Code runs it cost about 0.45 times as much per changed line. Agents re-send their whole context on every call, so fewer calls and fewer review rounds outweighed the higher token price.
Reality
- Evidence45
- Adoption8
- Hype gap+15
- Incentives
- Insufficient
- Confidence38
Cursor's Rollouts bot follows a pull request into production and grades the deploy against a plan it wrote when the PR opened. Its third verdict, inconclusive, is what you get when the signals were never instrumented.
Reality
- Evidence48
- Adoption22
- Hype gap+18
- Incentives74
- Confidence46
A dev.to checklist of ten software delivery gates for AI agents insists every check be automated, fast, blocking and observable. The gate its author calls no replacement for human review cannot meet that bar.
Reality
- Evidence22
- Adoption
- Insufficient
- Hype gap+40
- Incentives20
- Confidence58
Boris Cherny named six guardrails for Claude-written production code at Anthropic. Almost everything else in circulation about how they fit together comes from the dev.to post that relayed his quote.
Reality
- Evidence25
- Adoption20
- Hype gap+45
- Incentives55
- Confidence60
Every plugin and theme release now sits in a six-hour cooldown while AI models and Jetpack Scan grade the changes, and a high enough risk score halts distribution before any human looks at it. WordPress says the check already caught a backdoor.
Reality
- Evidence45
- Adoption68
- Hype gap+18
- Incentives62
- Confidence55
The WordPress.org update API will refuse any plugin release its own scoring calls high risk. The decision moves off a team member's inbox and onto a threshold the announcement does not publish.
Reality
- Evidence55
- Adoption68
- Hype gap+15
- Incentives60
- Confidence58
Datadog Security Labs put one document-portal prompt through three coding agents in both modes and audited all six builds. An insecure direct object reference appeared in each, and one build per cell leaves mode effects entangled with noise.
Publishers:securitylabs.datadoghq.com
Reality
- Evidence52
- Adoption
- Insufficient
- Hype gap+12
- Incentives68
- Confidence45
A dev.to tip argues agent pipelines need an adversarial verifier from a different model family plus randomized human audits, because same-model review exploits a documented self-preference bias.
Reality
- Evidence20
- Adoption
- Insufficient
- Hype gap+38
- Incentives
- Insufficient
- Confidence30