On PR #149, Codex filed 68 comments across 10 rounds and the team fixed 62 of them, all correct. The write-up blames the house rule that turned every comment into a commit for the code that came out worse.
Reality
- Evidence48
- Adoption20
- Hype gap+10
- Incentives40
- Confidence55
A team pushing 70 to 100 pull requests a week reports a 30 to 40 percent first-pass catch rate across 4,200 PRs. The parts that travel are the file-level chunk boundary and the draft gate; the percentage stays with the stack.
Reality
- Evidence38
- Adoption18
- Hype gap+32
- Incentives40
- Confidence45
Code review has always been sampling, and the signals it sampled stood in for a mind that had read the code. A dev.to post argues that when diffs arrive without one, what needs rewriting is the process, not the effort.
Reality
- Evidence32
- Adoption
- Insufficient
- Hype gap+14
- Incentives22
- Confidence48
Canvas arc() wants radians while CSS transform wants degrees, and the mismatch never raises an exception, so a dev.to checklist puts the catch in the diff and the lint config instead of in each bug fix.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+12
- Incentives35
- Confidence45
A 90-day instrumented study reports 28% more pull requests and 13 points of test coverage, but the gain lands only after training and workflow embedding, which puts the money on enablement rather than seats.
Reality
- Evidence45
- Adoption38
- Hype gap+12
- Incentives
- Insufficient
- Confidence40
Bacchelli and Bird classified 570 review comments inside Microsoft: 14 percent were about defects, 29 percent about improving code. Most review scoreboards still count bugs.
Reality
- Evidence44
- Adoption
- Insufficient
- Hype gap+30
- Incentives26
- Confidence41
A practitioner argues the 80% coverage norm was a budget for review time, not a standard, and that any check living only in the pipeline reports to you rather than constraining the agent.
Reality
- Evidence20
- Adoption14
- Hype gap+28
- Incentives
- Insufficient
- Confidence34
A published Playwright and Cucumber review checklist puts flakiness at the merge rather than the runner: state isolation, synchronisation strategy and single-behaviour scoping.
Reality
- Evidence30
- Adoption10
- Hype gap+25
- Incentives30
- Confidence52