Anthropic released Sonnet 5.5 at $2 and $10 per million input and output tokens, half the Opus 5.5 rate. How much a buyer saves by moving work down a tier depends on tokens burned per task and on cache reads priced identically on both models.
Perspective Coverage
5 publishers
- Builder
- Builder 46%
- Operator
- Operator 38%
- Investor
- Investor 16%
Reality
- Evidence55
- Adoption35
- Hype gap+20
- Incentives70
- Confidence60
AI code reviewers lose hosted rule ingestion first on self-managed GitLab, Azure DevOps Server and Bitbucket Data Center, a dev.to comparison says. Of the three mechanisms it compares, only a pass/fail rule can block a merge, and vendors tend to put that one on paid tiers.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+20
- Incentives
- Insufficient
- Confidence40
LinearB's 2026 benchmark finds only 32.7% of AI-assisted pull requests accepted within 30 days, against 84.4% of manual ones. The figures are associations, but they put review cost in the queue and in repeat rounds, where faster diff reading helps little.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+10
- Incentives40
- Confidence40
TypeSafe AI's Jev, a model that returns scored decisions, is in some developers' products after a launch video drew almost 40 million views on X. The evidence so far is hackathon testimony, so the useful work for leaders this quarter is finding which routine decisions can be split into narrow questions.
Reality
- Evidence30
- Adoption20
- Hype gap+45
- Incentives60
- Confidence35
Copilot code review for Azure Repos needs admins at three levels to enable it, in order, before it can comment on one pull request. It is a limited preview with no SLA, so whether review runs at all is configuration an Azure DevOps shop has to own and keep checking.
Reality
- Evidence50
- Adoption
- Insufficient
- Hype gap0
- Incentives35
- Confidence50
WorldScript Studio tracks fifteen automated reviewers in a JSON registry, and only four deterministic security scanners may block a merge. The design keeps LLM false positives off the merge path and keeps pull-request code away from the checker that judges it.
Reality
- Evidence50
- Adoption
- Insufficient
- Hype gap+10
- Incentives55
- Confidence45
LinearB scored 16 hand-built bugs on noise and clarity; DeepSource ran the same product category against OpenSSF's public CVE corpus and published F1. Each vendor wins on its own instrument, and the two tests reward different behaviour.
Reality
- Evidence42
- Adoption
- Insufficient
- Hype gap+38
- Incentives86
- Confidence52
Code Review Bench reconstructs the timeline of 16,017 open source pull requests and publishes precision and recall beside every F1. The top five tools sit within five points of each other, on samples of very different size.
Reality
- Evidence66
- Adoption45
- Hype gap+12
- Incentives55
- Confidence55
The case that generated code fails without the usual warning signs comes from two practitioner accounts, while the only measured series is curl's inbound bug reports, and Daniel Stenberg stopped taking those in January 2026.
Reality
- Evidence38
- Adoption58
- Hype gap+34
- Incentives62
- Confidence46
Crunchbase counted 29 additions to its unicorn board in August, worth about $63bn, and across the nine AI, software and semiconductor entrants with disclosed terms, every dollar of new cash set roughly eight dollars of paper.
Reality
- Evidence50
- Adoption24
- Hype gap+32
- Incentives66
- Confidence45
A dev.to post declines to rank the two on review accuracy without a controlled head-to-head, and asks instead whether you want a contract that ends in a tested working tree or feedback that stays attached to the PR.
Reality
- Evidence38
- Adoption
- Insufficient
- Hype gap−8
- Incentives25
- Confidence45
A retry loop with backoff, a timeout and test coverage charged 1,100 customers twice because it read an HTTP timeout as proof nothing happened. The diff never said otherwise.
Reality
- Evidence24
- Adoption32
- Hype gap+34
- Incentives44
- Confidence30
GitHub's own docs say Copilot always submits a Comment review. A governed change has four checkpoints, and the one holding authority is still branch protection.
Reality
- Evidence46
- Adoption
- Insufficient
- Hype gap−8
- Incentives54
- Confidence42
AI is billed by usage, and the introductory pricing has ended. The programs in trouble are the ones that put exhaustive, repeatable checking work onto a metered model endpoint.
Reality
- Evidence30
- Adoption34
- Hype gap+32
- Incentives84
- Confidence33
Tessl Code Review, free during beta, keeps review criteria as versioned files the team owns instead of logic sealed inside a vendor product. Someone still has to write the criteria.
Reality
- Evidence28
- Adoption10
- Hype gap+38
- Incentives68
- Confidence44
Microsoft has delayed the first cumulative update for Exchange Server Subscription Edition a second time, with no new date, while engineers validate a growing pile of AI-found security findings.
Reality
- Evidence58
- Adoption24
- Hype gap+22
- Incentives62
- Confidence55
Reported figures put bug rates 41% higher and security flaws 2.74x more prevalent in AI assisted code. The constraint that decides whether that matters is review throughput.
Reality
- Evidence24
- Adoption31
- Hype gap+34
- Incentives58
- Confidence21
A vendor essay on AI package hallucination makes a defensible case: a package name that does not exist yet cannot be scanned, so the control has to sit at selection.
Reality
- Evidence24
- Adoption22
- Hype gap+41
- Incentives88
- Confidence45