Skip to content

Topic

AI Code Review

The practice of using large language models or AI agents to review source code and pull requests for bugs, style issues, and security flaws.

Current stories

build1 publisher

Idiomatic os.path.join trips Gemini 3.7 Flash in a 12-task LLM security benchmark

Six LLMs on a 12-task Kaggle security benchmark all caught SQL injection, hardcoded keys and pickle RCE, but Gemini 3.7 Flash missed a path traversal. With one scenario per flaw class, the run shows which textbook patterns the models know and says little about trusting one to review real code.

Publishers:dev.to

Reality

Evidence30
Adoption
Insufficient
Hype gap+40
Incentives40
Confidence35
build1 publisher

GitLab gives Duo Enterprise seat holders the single-pass AI reviewer by default

GitLab's v19.5 docs send AI reviews started by Duo Enterprise seat holders to the single-pass reviewer that reads only the merge request and its diffs. To check cross-file rules on every review, a group Owner has to move seat holders onto the credit-billed Code Review Flow.

Publishers:dev.to

Reality

Evidence45
Adoption
Insufficient
Hype gap+5
Incentives
Insufficient
Confidence40
build1 publisher

A 76,326-word BMAD plan blocked its own implementation on day five

One operator ran spec-driven development at full BMAD weight for four weeks on an autonomous coding repo. The planning stage produced 1,782 files and 16 MB of artifacts, and a single story burned around thirty million tokens.

Publishers:dev.to

Reality

Evidence34
Adoption20
Hype gap+20
Incentives45
Confidence38
build1 publisher

The 90-second diff still costs 20 minutes of human review

The expensive part of software work has moved from writing code to proving it. An engineer building agent systems puts the deny outside the model where no prompt can reach it, and says plainly which defects his deterministic checks still miss.

Publishers:dev.to

Reality

Evidence36
Adoption14
Hype gap+22
Incentives52
Confidence44

Earlier coverage

  1. A four-hour Claude Code session landed 42 unreviewed commits on main

    Build · September 14, 2026 · 1 publisher

  2. Fixing 62 true review comments left AgentCoop guarding a config key nobody writes

    Build · September 14, 2026 · 1 publisher

  3. Since 1 September a Copilot approval can satisfy GitHub's required-approvals rule

    Build · September 13, 2026 · 1 publisher

  4. A proposed circuit breaker for code-review agents defaults its security guardrail to None

    Build · September 11, 2026 · 1 publisher

  5. One prompt turned 46 SWC issues into 45 parallel pull requests on one subsystem

    Build · September 11, 2026 · 1 publisher

  6. Second Codex pass on the fixed pull request found four more distinct defects

    Build · September 10, 2026 · 1 publisher

  7. Control for reasoning before you credit spec-first prompting for the quality gain

    Build · September 10, 2026 · 1 publisher

  8. OpenAI hands merge-blocking authority to its own security review model

    Build · September 9, 2026 · 1 publisher

  9. A six-person team's weekly review load grew 7.6x in a quarter

    Build · September 8, 2026 · 1 publisher

  10. Reviewing every PR with Claude first changes what the senior reviewer looks for

    Build · September 8, 2026 · 1 publisher

  11. Requiring a reproduction path stops a review agent from quadrupling the feature

    Build · September 6, 2026 · 1 publisher

  12. Reviewing agent diffs by taste costs you the one finding that pages someone

    Build · September 2, 2026 · 1 publisher

  13. GitHub lets Copilot's sign-off count toward a repository's required approvals

    Product · September 2, 2026 · 1 publisher

  14. Make agent-written pull requests carry a receipt

    Build · September 2, 2026 · 1 publisher

  15. Two synthetic PRs measure whether ADR-0012 outranks the reviewer's own memory

    Build · August 31, 2026 · 1 publisher

  16. A code graph caught the login bug sitting three hops and an event bus from the diff

    Build · August 30, 2026 · 1 publisher

  17. An approving LLM comment sent an unguarded array index into a payment reconciliation job

    Build · August 29, 2026 · 1 publisher

  18. A client-set header decided whether a Copilot request spent premium quota

    Build · August 28, 2026 · 1 publisher

  19. A decision rule sorts Codex and CodeRabbit by the unit of work each owns to completion

    Build · August 28, 2026 · 1 publisher

  20. Uber keeps its AI bill flat by choosing which model runs which workload

    Leadership · August 28, 2026 · 1 publisher

  21. Harness gives the coding agent its own permissions and its own audit trail

    Product · August 27, 2026 · 1 publisher

  22. The bug that was not in the diff: 1,100 double charges cleared a four-minute review

    Build · August 26, 2026 · 1 publisher

  23. Copilot code review only ever files comments, so stop buying it as a merge gate

    Build · August 25, 2026 · 1 publisher

  24. Copilot's meter changed on June 1, and half your seats are still priced in the old unit

    Build · August 25, 2026 · 1 publisher

  25. LinkedIn graded its own AI reviewer against merged code, and 63.9% of comments stuck

    Build · August 22, 2026 · 1 publisher

  26. Tessl moves review standards into the repo, and hands teams the homework

    Product · August 20, 2026 · 1 publisher

  27. An AI reviewer called injectable SQL safe because it could not read the helper

    Build · August 19, 2026 · 1 publisher

  28. Code review was the apprenticeship, and AI diffs are ending it without a replacement

    Build · August 19, 2026 · 1 publisher

  29. Opus 5 absorbed your verify prompts. The reading is still on your desk.

    Build · August 19, 2026 · 1 publisher

  30. Fuse ranks, not scores: a retrieval contract that refuses to guess in code review

    Build · August 17, 2026 · 1 publisher

  31. Your reviewing model is reading the diff when it should be reading the session

    Build · August 14, 2026 · 1 publisher