Build1 publisher3 min readPublished
Code review was the apprenticeship, and AI diffs are ending it without a replacement
A New Stack column argues AI-generated diffs have made careful review impossible. The automatable half of the job is bug-catching; the half nobody has staffed is knowledge transfer.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- The volume of AI-generated code has broken code review: engineers are reviewing 500-line diffs they did not write, generated by models they do not fully control, at a pace that makes careful reading impossible.
- Code review has always been where knowledge moves through the team, where junior engineers watch how senior engineers think, where architectural decisions are challenged, mental models of the codebase are formed, and where shared ownership takes shape.
- Most conversation about replacing code review focuses on catching bugs and standards checking, and most of that work can be automated by LLM-based adversarial review, automated verification and deterministic checks.
- If standards-checking is solved completely and nothing is done about knowledge sharing, only the easier job has been fixed.
- In a pre-AI team, a junior's PR generates four to eight comments from a senior on idiomatic patterns, a back-and-forth on edge cases, and an implicit lesson in how the senior would think about the problem.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
A column published by The New Stack argues that the volume of AI-generated code has broken code review: engineers are reading 500-line diffs they did not write, generated by models they do not fully control, at a pace that makes careful reading impossible [1]. The second casualty is the one worth planning around, because review was also where knowledge moved through a team, where junior engineers watched senior engineers think, architectural decisions got challenged, and mental models of the codebase formed [2].
The defect half of the job is the tractable half. The piece's own claim is that most bug catching and standards checking can be handled by LLM-based adversarial review, automated verification and deterministic checks [3], and that solving that completely while doing nothing about knowledge sharing fixes only the easier job [4]. The teaching half never had tooling because it never needed any. It was a byproduct of somebody reading carefully with their name on the approval.
The before-and-after in the column is specific about what disappears. In a pre-AI team, a junior's pull request draws four to eight comments from a senior on idiomatic patterns, plus a back-and-forth on edge cases and an implicit lesson in how to think about the problem [5]. In an AI-heavy team, the same change is generated by AI, lightly edited, reviewed by AI, lightly approved, and merged with no human having formed a mental model of it [6]. Both failures have one cause: nobody read it [19].
Margaret-Anne Storey's name for the residue is cognitive debt. Technical debt lives in the code; cognitive debt lives in people, accumulating when a team stops understanding the system it is building, and no dashboard reports it [7][8]. Her evidence, as relayed in the piece, is observational rather than measured: a group of students building with AI eventually told her they could no longer make changes to their product [9]. She suspected messy code. What she found was that they had lost track of which features they were building and why, and did not know who on the team knew what; on one team a single person understood the code because they were supervising the generation, the rest could not, and the person who generated it did not really understand the output either [10][11].
The obvious rebuttal is to ask the model. The column concedes that LLMs explain how a feature or behavior works, and argues that what goes missing is the what and the why: why the change was built this way, what assumptions were made, what the side effects are, why the feature needs to exist [12][13]. None of that is wired into the code, so no reader of the code recovers it.
The proposed fix moves the human checkpoint upstream, to reviewing intent (plans, constraints and acceptance criteria) instead of the diff [14], and treats knowledge sharing the same way: capture the architectural choices, behavior tradeoffs and scope calls an engineer makes during a session with an agent, and write them down as acceptance criteria [15]. ThoughtWorks Market Technology Director Vanitha Kumar, in the podcast conversation the column draws on, built something adjacent: an agent instructed to check code against her team's reference practices, identify deviations and explain why each deviation would be detrimental, built as a teaching agent before it drifted into a review agent [16][17]. "If code review shifts left, knowledge sharing also has to shift left," the column argues [18].
Watch whether those criteria get written to be read. An acceptance-criteria file that only its author opens is a commit message with better formatting. The cheap test: take a subsystem your agents shipped last quarter and ask two engineers who did not supervise it to explain why it works the way it does.