Build1 publisherNot yet confirmed elsewhere3 min readPublished
AI's damage shows up at review, and a single approver has nothing downstream to catch it
A dev.to writer says he approved code he could not explain. The numbers around him point at the same place: the failure is in review, where one approver is the whole gate.
The Engineer · Build desk

What happened
- A dev.to author says he approved most of the code in last year's projects and could not explain a good part of it; tests were green and the diff looked familiar.
- The projects were his own, so no reviewer sat downstream of him and his approval was the only quality check in the path.
- In the 2025 Stack Overflow survey, one in five developers named lost confidence in their own problem-solving as a top frustration with AI tools.
- Casner's 2014 pilot study found hand-flying skills largely intact after years of automation while situational awareness and failure handling had collapsed.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- exposure The person who signs off on a merge is now potentially the same person who cannot reconstruct why the code does what it does, and the signature is what the org treats as assurance.
- constraint Retrospectives and self-reports cannot be used to measure this, because experienced developers misjudged their own throughput by tens of points on code they knew better than anyone.
- decision Teams that adopted one-approval branch rules to keep merges moving now have to decide whether that reader is a gate or a formality, because green CI no longer distinguishes the two.
- capability Anyone wanting a warning signal has to build it outside the individual, since the described mechanism removes the sensation of having stopped.
An approval leaves the same trace whether or not anyone read the diff. CI stays green, the branch merges, the queue moves. Nothing in the pipeline separates a change that was understood from one that merely looked like something the reviewer would have written himself [1], so approval latency and pass rate hold steady while the thing they stand in for drains away.
The loss becomes visible only at the second reader. With two reviewers, an unread approval is one layer of a two-layer catch, and the second layer is where it surfaces. The dev.to account is the other configuration: his own projects, nobody downstream of him, one gate, and by his description the gate had stopped reading [2].
METR's trial is the strongest test of self-knowledge in this material. Sixteen experienced developers, 246 real tasks, repositories carrying an average of five years of their own contributions [7]. Afterwards they put themselves at roughly 20 percent faster; the measurement said 19 percent slower, a 39-point miss on work they had just finished [8][10]. That is not a wrong opinion about a tool.
The aviation research explains why the damage lands on review rather than on authorship. Casner and colleagues found that hand-flying survived automation, and what cratered were the cognitive skills: knowing where the aircraft actually was, what the next step should be, what a failing instrument looked like [4]. Macnamara's 2024 review argues the AI case should be worse, because AI substitutes for cognitive work and cognitive skills decay faster than physical ones [6]. It also names why nobody catches it in themselves: ordinary decay is noticed because you stopped doing the task, and this decay hides because you kept shipping and only the engagement stopped [9].
What none of this material contains is a defect rate. The Stack Overflow figure counts developers reporting lost confidence in their own problem-solving [3], not bugs that escaped. The r/ExperiencedDevs thread counts agreement, four hundred replies to one person calling himself a tourist in his own codebase [12]. METR measured time, not review quality [8]. The author is straight about where that leaves the argument: the diagnosis is done, and every remedy anyone proposes has the same shape, a shape known to fail, including the one he spent two weeks designing [11].
The one intervention in the record aimed at automation decay rather than at tooling is the FAA's, which was to recommend pilots hand-fly the majority of flights [5]. It works by spending the automation's benefit back on purpose. A review process that wants the same effect has to buy the reps the same way, and pay for them out of the same budget the assistant was supposed to fill.
What to watch
- Whether METR's 19 percent slowdown replicates at larger n and outside mature open-source repositories.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence38
- Adoption
- Insufficient
- Hype gap+22
- Incentives45
- Confidence40
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
The dev.to author writes that he approved most of the code and commits that went into apps he built last year, and could not explain a good portion of it even to himself; the tests were green and the diff looked like something he would have written, so he hit approve.
- [2]
Because most of the projects were his own, there was no second reviewer downstream of him; he says he was the entire quality gate and the gate had quietly stopped reading.
- [3]
One in five developers in the 2025 Stack Overflow survey picked "I've become less confident in my own problem-solving" as a top frustration with AI tools.
- [4]
Casner et al., 2014 tested airline pilots who trained on manual flight but spent their careers flying highly automated aircraft: procedural skills such as scanning instruments and hand-flying were rusty but largely intact, while cognitive skills had cratered, including maintaining awareness of where the aircraft actually was, tracking what the next step should be, and recognising and handling an instrument failure.
- [5]
The FAA's response to the Casner findings was to recommend that pilots hand-fly for the majority of flights.
- [6]
A 2024 review by Brooke Macnamara and colleagues in Cognitive Research argues that AI-induced skill decay should be worse than the automation decay documented in cockpits, because AI mimics cognitive rather than mechanical work and cognitive skills decay faster than physical ones.
ReportedSupportedSource: Macnamara et al., 2024, Cognitive Research, as cited by dev.toView cited source - [7]
METR ran a randomised controlled trial in 2025 with sixteen experienced open-source developers on 246 real tasks in their own repositories, projects averaging five years of their own contributions.
- [8]
In the METR trial, developers predicted AI would make them 24% faster, afterwards estimated they had been about 20% faster, and were measured 19% slower.
- [9]
Macnamara's review names the mechanism by which AI-induced decay hides: ordinary skill decay is noticed because the person stopped doing the task, whereas AI-induced decay leaves the task in place (shipping, reviewing, closing tickets) and only cognitive engagement stops, so a surgeon still completing successful operations has no signal that judgment has softened.
- [10]
The gap between the METR participants' post-hoc self-estimate and the measured result is 39 percentage points; against their prior prediction the gap is 43 points.
- [11]
The author says the diagnosis is already done and that every fix anyone proposes for the problem has the same shape, a shape known to fail, including the fix he spent two weeks designing.
ReportedInsufficientSource: dev.to, Michael2 sources— create a free account to open themView cited source - [12]
The author cites an r/ExperiencedDevs thread in which someone calls themselves "a tourist in their own codebase" and four hundred people reply in agreement.
Sources
1 independent publisher whose own reporting we read for this story.
- dev.toAI didn't make me a worse coder. It made me a worse reviewer.
1 article · August 21, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.
Topics
- Code Review and Approval GatesFollow
- Automation-Induced Skill DecayFollow
- AI-Assisted CodingFollow
- Metacognition and Self-AssessmentFollow
- Developer Productivity MeasurementFollow