GitHub's September 23 Copilot code review update adds separate automatic-review switches for draft pull requests and new pushes, plus Lite or Balanced effort. Its review is still a comment that cannot satisfy a merge rule, so each extra trigger adds reading without adding a gate.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap−10
- Incentives30
- Confidence45
Six LLMs on a 12-task Kaggle security benchmark all caught SQL injection, hardcoded keys and pickle RCE, but Gemini 3.7 Flash missed a path traversal. With one scenario per flaw class, the run shows which textbook patterns the models know and says little about trusting one to review real code.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+40
- Incentives40
- Confidence35
AI code reviewers lose hosted rule ingestion first on self-managed GitLab, Azure DevOps Server and Bitbucket Data Center, a dev.to comparison says. Of the three mechanisms it compares, only a pass/fail rule can block a merge, and vendors tend to put that one on paid tiers.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+20
- Incentives
- Insufficient
- Confidence40
LinearB's 2026 benchmark finds only 32.7% of AI-assisted pull requests accepted within 30 days, against 84.4% of manual ones. The figures are associations, but they put review cost in the queue and in repeat rounds, where faster diff reading helps little.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+10
- Incentives40
- Confidence40
GitLab's v19.5 docs send AI reviews started by Duo Enterprise seat holders to the single-pass reviewer that reads only the merge request and its diffs. To check cross-file rules on every review, a group Owner has to move seat holders onto the credit-billed Code Review Flow.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+5
- Incentives
- Insufficient
- Confidence40
Copilot code review for Azure Repos needs admins at three levels to enable it, in order, before it can comment on one pull request. It is a limited preview with no SLA, so whether review runs at all is configuration an Azure DevOps shop has to own and keep checking.
Reality
- Evidence50
- Adoption
- Insufficient
- Hype gap0
- Incentives35
- Confidence50
WorldScript Studio tracks fifteen automated reviewers in a JSON registry, and only four deterministic security scanners may block a merge. The design keeps LLM false positives off the merge path and keeps pull-request code away from the checker that judges it.
Reality
- Evidence50
- Adoption
- Insufficient
- Hype gap+10
- Incentives55
- Confidence45
GitHub's agentic autofix now saves a pattern from each security fix it creates in Copilot Memory and passes those patterns to code review and the cloud agent. Teams that enable Memory now decide whether an agent's fixes deserve to guide reviews across a repository.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+15
- Incentives45
- Confidence50
Blog vs Bytecode, a 28-item Kaggle benchmark, graded empty proxy responses as wrong and scored DeepSeek-R1 at 17% until a second gateway showed 100%. Once capture was fixed, frontier models lost points by flagging sound code, while a small Gemma model missed most of the planted flaws.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+10
- Incentives30
- Confidence45
LinearB scored 16 hand-built bugs on noise and clarity; DeepSource ran the same product category against OpenSSF's public CVE corpus and published F1. Each vendor wins on its own instrument, and the two tests reward different behaviour.
Reality
- Evidence42
- Adoption
- Insufficient
- Hype gap+38
- Incentives86
- Confidence52
One operator ran spec-driven development at full BMAD weight for four weeks on an autonomous coding repo. The planning stage produced 1,782 files and 16 MB of artifacts, and a single story burned around thirty million tokens.
Reality
- Evidence34
- Adoption20
- Hype gap+20
- Incentives45
- Confidence38
Code Review Bench reconstructs the timeline of 16,017 open source pull requests and publishes precision and recall beside every F1. The top five tools sit within five points of each other, on samples of very different size.
Reality
- Evidence66
- Adoption45
- Hype gap+12
- Incentives55
- Confidence55
A two-model code review produced rebuttals and a clean verdict while the raw logs showed no changed positions and no new evidence. The rebuild enforces independence by withholding each verdict until both are committed.
Reality
- Evidence34
- Adoption10
- Hype gap+22
- Incentives48
- Confidence44
GitHub's podcast companion post argues that a production authentication refactor and a CSS experiment need different review, and it puts the same conditional test to "RAG is dead" and "Skills killed MCP".
Reality
- Evidence44
- Adoption
- Insufficient
- Hype gap+6
- Incentives58
- Confidence63
Qodo's five-agent comparison ranks Qodo first and publishes no method behind the grades, while Graphite's own integration guide lists four ways to wire a reviewer into GitHub and only one of them owns merge protection.
Reality
- Evidence42
- Adoption
- Insufficient
- Hype gap+32
- Incentives70
- Confidence46
The New Stack says AI-assisted teams have moved the bottleneck from writing code to verifying it, and the remedy it proposes is an afternoon spent sorting the last 100 review comments into rules, tests and judgment.
Reality
- Evidence32
- Adoption15
- Hype gap+35
- Incentives80
- Confidence55
Severity labels come back from Claude, get checked against four allowed values, stable-sorted and then cut with slice() at a constant, with the count of what got cut printed in the review body, while the priority order inside the cut still comes from the model.
Reality
- Evidence64
- Adoption12
- Hype gap−8
- Incentives58
- Confidence55
The expensive part of software work has moved from writing code to proving it. An engineer building agent systems puts the deny outside the model where no prompt can reach it, and says plainly which defects his deterministic checks still miss.
Reality
- Evidence36
- Adoption14
- Hype gap+22
- Incentives52
- Confidence44
Alibaba ran the reviewer internally for two years across tens of thousands of developers before releasing it. The write-up is specific about how files get selected and comments get positioned, and it stops short of publishing the benchmark scores.
Reality
- Evidence32
- Adoption30
- Hype gap+55
- Incentives68
- Confidence38
In an InfoQ talk, Duolingo engineer Sarah Deitke describes the hands-on labs and usage dashboards her DevEx AI team built to make engineers comfortable with AI, and says they trust in-house content over a vendor's.
Reality
- Evidence45
- Adoption25
- Hype gap+30
- Incentives55
- Confidence55
Earlier coverage
- A four-hour Claude Code session landed 42 unreviewed commits on main
Build · September 14, 2026 · 1 publisher
- Fixing 62 true review comments left AgentCoop guarding a config key nobody writes
Build · September 14, 2026 · 1 publisher
- Since 1 September a Copilot approval can satisfy GitHub's required-approvals rule
Build · September 13, 2026 · 1 publisher
- A proposed circuit breaker for code-review agents defaults its security guardrail to None
Build · September 11, 2026 · 1 publisher
- One prompt turned 46 SWC issues into 45 parallel pull requests on one subsystem
Build · September 11, 2026 · 1 publisher
- Second Codex pass on the fixed pull request found four more distinct defects
Build · September 10, 2026 · 1 publisher
- Control for reasoning before you credit spec-first prompting for the quality gain
Build · September 10, 2026 · 1 publisher
- OpenAI hands merge-blocking authority to its own security review model
Build · September 9, 2026 · 1 publisher
- A six-person team's weekly review load grew 7.6x in a quarter
Build · September 8, 2026 · 1 publisher
- Reviewing every PR with Claude first changes what the senior reviewer looks for
Build · September 8, 2026 · 1 publisher
- Requiring a reproduction path stops a review agent from quadrupling the feature
Build · September 6, 2026 · 1 publisher
- Reviewing agent diffs by taste costs you the one finding that pages someone
Build · September 2, 2026 · 1 publisher
- GitHub lets Copilot's sign-off count toward a repository's required approvals
Product · September 2, 2026 · 1 publisher
- Make agent-written pull requests carry a receipt
Build · September 2, 2026 · 1 publisher
- Two synthetic PRs measure whether ADR-0012 outranks the reviewer's own memory
Build · August 31, 2026 · 1 publisher
- A code graph caught the login bug sitting three hops and an event bus from the diff
Build · August 30, 2026 · 1 publisher
- An approving LLM comment sent an unguarded array index into a payment reconciliation job
Build · August 29, 2026 · 1 publisher
- A client-set header decided whether a Copilot request spent premium quota
Build · August 28, 2026 · 1 publisher
- A decision rule sorts Codex and CodeRabbit by the unit of work each owns to completion
Build · August 28, 2026 · 1 publisher
- Uber keeps its AI bill flat by choosing which model runs which workload
Leadership · August 28, 2026 · 1 publisher
- Harness gives the coding agent its own permissions and its own audit trail
Product · August 27, 2026 · 1 publisher
- The bug that was not in the diff: 1,100 double charges cleared a four-minute review
Build · August 26, 2026 · 1 publisher
- Copilot code review only ever files comments, so stop buying it as a merge gate
Build · August 25, 2026 · 1 publisher
- Copilot's meter changed on June 1, and half your seats are still priced in the old unit
Build · August 25, 2026 · 1 publisher
- LinkedIn graded its own AI reviewer against merged code, and 63.9% of comments stuck
Build · August 22, 2026 · 1 publisher
- Tessl moves review standards into the repo, and hands teams the homework
Product · August 20, 2026 · 1 publisher
- An AI reviewer called injectable SQL safe because it could not read the helper
Build · August 19, 2026 · 1 publisher
- Code review was the apprenticeship, and AI diffs are ending it without a replacement
Build · August 19, 2026 · 1 publisher
- Opus 5 absorbed your verify prompts. The reading is still on your desk.
Build · August 19, 2026 · 1 publisher
- Fuse ranks, not scores: a retrieval contract that refuses to guess in code review
Build · August 17, 2026 · 1 publisher
- Your reviewing model is reading the diff when it should be reading the session
Build · August 14, 2026 · 1 publisher