Build1 distinct publisher3 min readUpdated
A dev.to writer says he approved code he could not explain. The numbers around him point at the same place: the failure is in review, where one approver is the whole gate.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
An approval leaves the same trace whether or not anyone read the diff. CI stays green, the branch merges, the queue moves. Nothing in the pipeline separates a change that was understood from one that merely looked like something the reviewer would have written himself [1], so approval latency and pass rate hold steady while the thing they stand in for drains away.
The loss becomes visible only at the second reader. With two reviewers, an unread approval is one layer of a two-layer catch, and the second layer is where it surfaces. The dev.to account is the other configuration: his own projects, nobody downstream of him, one gate, and by his description the gate had stopped reading [2].
METR's trial is the strongest test of self-knowledge in this material. Sixteen experienced developers, 246 real tasks, repositories carrying an average of five years of their own contributions [8]. Afterwards they put themselves at roughly 20 percent faster; the measurement said 19 percent slower, a 39-point miss on work they had just finished [9][10]. That is not a wrong opinion about a tool.
The aviation research explains why the damage lands on review rather than on authorship. Casner and colleagues found that hand-flying survived automation, and what cratered were the cognitive skills: knowing where the aircraft actually was, what the next step should be, what a failing instrument looked like [5]. Macnamara's 2024 review argues the AI case should be worse, because AI substitutes for cognitive work and cognitive skills decay faster than physical ones [7]. It also names why nobody catches it in themselves: ordinary decay is noticed because you stopped doing the task, and this decay hides because you kept shipping and only the engagement stopped [11].
What none of this material contains is a defect rate. The Stack Overflow figure counts developers reporting lost confidence in their own problem-solving [4], not bugs that escaped. The r/ExperiencedDevs thread counts agreement, four hundred replies to one person calling himself a tourist in his own codebase [3]. METR measured time, not review quality [9]. The author is straight about where that leaves the argument: the diagnosis is done, and every remedy anyone proposes has the same shape, a shape known to fail, including the one he spent two weeks designing [12].
The one intervention in the record aimed at automation decay rather than at tooling is the FAA's, which was to recommend pilots hand-fly the majority of flights [6]. It works by spending the automation's benefit back on purpose. A review process that wants the same effect has to buy the reps the same way, and pay for them out of the same budget the assistant was supposed to fill.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
The dev.to author writes that he approved most of the code and commits that went into apps he built last year, and could not explain a good portion of it even to himself; the tests were green and the diff looked like something he would have written, so he hit approve.
Because most of the projects were his own, there was no second reviewer downstream of him; he says he was the entire quality gate and the gate had quietly stopped reading.
One in five developers in the 2025 Stack Overflow survey picked "I've become less confident in my own problem-solving" as a top frustration with AI tools.
Casner et al., 2014 tested airline pilots who trained on manual flight but spent their careers flying highly automated aircraft: procedural skills such as scanning instruments and hand-flying were rusty but largely intact, while cognitive skills had cratered, including maintaining awareness of where the aircraft actually was, tracking what the next step should be, and recognising and handling an instrument failure.
The FAA's response to the Casner findings was to recommend that pilots hand-fly for the majority of flights.
A 2024 review by Brooke Macnamara and colleagues in Cognitive Research argues that AI-induced skill decay should be worse than the automation decay documented in cockpits, because AI mimics cognitive rather than mechanical work and cognitive skills decay faster than physical ones.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Single practitioner essay relaying uncorroborated external research
The cluster contains exactly one source: a first-person dev.to essay. Its strongest evidence is secondhand — a relayed 2025 METR randomised trial (n=16, 246 tasks), Casner et al. 2014, a 2024 Macnamara review, and a Stack Overflow survey figure — with no links, methodology detail, or independent confirmation available in the cluster. The self-report and the unlinked Reddit thread add colour but not verification, and the central prescriptive assertion is unenumerated.
No adoption signal in supplied sources
The cluster reports no release, deployment, pricing, licensing, benchmark run, or usage disclosure. It describes a behavioural and review-process hypothesis, and the supplied material contains no evidence of any practice, tool, or policy being adopted in response.
Modest overreach: universal claims on one anecdote plus a relayed n=16 trial
The essay is unusually self-limiting for the genre — it concedes the diagnosis is well-worn and offers a falsifiable self-test rather than a product — which keeps the gap small. It is still overstated relative to its evidence: a sweeping 'every proposed fix has the same shape and that shape is known to fail' is asserted without enumeration, prevalence rests on an unlinked thread, and a sixteen-developer relayed trial is treated as a general calibration failure for all developers.
Engagement-seeking developer-platform essay; no vendor stake disclosed
The author writes on a developer publishing platform and structures the piece for response — an explicit exercise, a request to report back, and an acknowledgement that this post format is already popular this spring — which rewards a resonant, alarming framing. No vendor, sponsor, employer, or commercial interest is disclosed or apparent, and the author volunteers that his own two-week fix failed, which cuts against a promotional motive.
Low: one publisher, secondhand figures, no adoption evidence
Attribution of what the article says is clear and quotable, so the claim set is reliable as reported speech. Confidence in the underlying facts is low: single publisher, every research figure relayed without primary reference, no adoption dimension measurable, and two claims rated insufficient on their own terms.
product
Engineering counts merged pull requests and nothing for the hours spent watching the agent1 distinct publisher
product
The 19% Gap: Why Developer Velocity Self-Reports Cannot Justify an AI Rollout1 distinct publisher
build
Six MariaDB versions, one real difference: the only reason to leave 10.6 is the July 2026 clock1 distinct publisher
build
Force the tool call, then hand Lightsail a long-lived key1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 21, 2026