Build1 distinct publisher2 min readPublished
Bacchelli and Bird classified 570 review comments inside Microsoft: 14 percent were about defects, 29 percent about improving code. Most review scoreboards still count bugs.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
The defect bucket in the Bacchelli and Bird sample holds roughly 80 comments, because 14 percent of 570 is 79.8, and the code-improvement bucket holds about 165 [1][3][4][1][2]. That is a thin evidential base for a claim as load-bearing as "review is not mainly bug-hunting," and worth saying before anyone rewrites a checklist around it. The later Microsoft paper puts possible-defect comments near 15 percent, within a point of the earlier figure [5][3], and reports at least half of comments concerned long-term maintainability [5] - more than three times the defect share [7]. Google's 2018 case study arrived somewhere similar by a different route, reading 12 interviews and 44 survey responses against logs from 9 million reviewed changes, and concluding that review served as teaching and codebase maintenance [6][7].
A comment classification measures what reviewers typed, not what they prevented. The 30-point distance between the 44 percent who rank defect-finding as their top motivation and the 14 percent of comments that were defect-related [2][3][5] is not a clean contradiction, and the dev.to writeup concedes the point: stated motivation and classified output are different instruments [10]. It is still an awkward gap. If a team's review metric is defects caught per review, the metric is tracking the smallest category anyone measured, and is silent on the largest.
The category no defect counter can see is the one the writeup is actually about: whether the change belongs where it was put [8]. Its worked example is a composite with invented details [9]. A payment path runs validate, authorize, capture, settle, and someone adds currency conversion inside authorize. The conversion rounds correctly and has unit tests. But the settlement amount is already computed upstream and shipped with the order, so the change couples the authorization path to an exchange-rate source it never needed, puts rounding and rate-selection policy in a second place where it can drift, and establishes that authorize may alter pricing, which opens the door to fees and discounts landing there next [9]. No syntax rule surfaces any of that, and more tests around the conversion do not address it either [11].
That is why the artifact of a good review is often a decision rather than a comment thread, which is the position the writeup ends on: fewer comments, better decisions [12]. A team that wanted to measure the work the research actually found would have to count something like reviews that changed where code lives, or that sent a responsibility back to its owner. Nobody's dashboard reports that. Meanwhile a review whose only output was a caught bug has, by omission, ratified every structural assumption in the diff [8].
Ranked by verification strength, evidence, and original report placement.
Finding defects was the top-ranked motivation for 44% of the developers surveyed in the Bacchelli and Bird study.
Only 14% of the comments in the Bacchelli and Bird sample were classified as defect-related.
A later Microsoft paper titled "Code Reviews Do Not Find Bugs" reported that about 15% of review comments indicated a possible defect, while at least 50% concerned long-term maintainability.
The dev.to writeup argues the highest-value review comment is "Should this code exist here at all?", because a change can be locally correct and systemically wrong, and such changes are dangerous precisely because they look good in the diff.
In a large Microsoft study, Alberto Bacchelli and Christian Bird observed 17 developers across 16 teams, manually classified 570 review comments, and surveyed 165 managers and 873 developers and testers.
Code improvements, such as removing unnecessary code or improving readability, were the largest comment category in the Bacchelli and Bird sample at 29%.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Named studies, cited second-hand
The descriptive core is anchored in three identifiable empirical sources with methodology detail (17 developers across 16 teams, 570 classified comments, 165 managers and 873 developers surveyed; a later Microsoft paper with ~15% defect and >=50% maintainability comments; a 2018 Google study spanning 12 interviews, 44 survey responses and 9 million reviewed changes). Against that, the cluster contains one secondary retelling with no citations, years or links for the first study, and the prescriptive half of the argument rests on a self-declared invented scenario plus untested assertions that linting and tests cannot surface placement problems.
No adoption signal supplied
The cluster contains no release, deployment, benchmark, usage disclosure or any account of a team adopting the described review practice or changing its review metrics. The cited studies describe existing review behaviour inside Microsoft and Google, not uptake of this article's prescription, so no adoption level can be measured without inferring facts the sources do not provide.
Descriptive numbers oversold as prescriptive proof
The cited percentages support only that review comments skew toward maintainability and improvement over defects. They say nothing about which comment type creates the most value, yet the framing presents them as backing the 'should this code exist here at all?' thesis. The article itself concedes stated motivation and comment classification are not the same measure, and its worked example is explicitly invented, so the strongest prescriptive claims outrun the evidence even though the underlying empirical summary is accurate and appropriately hedged.
Low commercial stake, authority-building platform
The single supplied source is a practitioner essay on a community publishing platform with no product, vendor, pricing or license interest anywhere in the text, and it cites third-party research it did not produce. The residual incentive visible in the material is reputational: a headline asserting what senior engineers 'actually' do and a provocative paper title used for emphasis, both of which reward engagement on that platform.
Single publisher, verifiable core, unmeasured prescription
Confidence is limited by a one-source, one-publisher cluster with no primary papers attached and no adoption evidence at all. It is supported by internal consistency: two independent Microsoft results agree within a point, the Google study points the same way, and the article discloses both its measurement caveat and the invented status of its example, so the descriptive claims can be held with moderate confidence while the prescriptive ones cannot.
build
Grok 4.6 lands in Copilot two days after launch, and the model picker becomes a procurement problem1 distinct publisher
product
Washington's secret AI test is coming for open weights, and release dates go with it2 distinct publishers
invest
Nvidia's August 26 print: 92% of the quarter rides on one segment1 distinct publisher
product
Cisco and Nvidia go looking for the other third of AI spending1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 26, 2026