Build1 publisher3 min readPublished
Opus 5 absorbed your verify prompts. The reading is still on your desk.
Anthropic says the model now verifies by default, so the double-check lines can go. The four-hundred-line diff behind the green check still has to be read by someone.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Anthropic's Opus 5 guidance says the model verifies by default, and the "verify" and "double-check" lines developers used to write now only make it over-verify.
- Checking its own work as it writes is one job and reviewing the finished diff is another; Opus 5 got better at the first, which quietly excuses the developer from the second.
- The four-hundred-line change the model produces lands on your branch behind a green check nobody actually read.
- A developer on Hacker News: "Once AI starts generating code, it flows out like water through a burst dam. It's impossible for any human to fully understand it all."
- Another Hacker News developer: "I also fire off tons of parallel agents, and review is hands down the biggest bottleneck."
Compiled by The EngineerSomething wrong?How this is made
Why it matters
Anthropic's Opus 5 guidance tells you to drop the "verify" and "double-check" lines from your prompts, because the model verifies by default and the extra instruction only makes it over-verify [1]. That deletes a prompting chore and leaves the expensive part of the pipeline exactly where it was: checking your work as you write it and reading a finished diff are two different jobs, and only the first one improved [2].
What lands on the branch is a four-hundred-line change behind a green check nobody actually read [3]. Anyone running several agents at once recognises the shape. One developer on Hacker News described the volume as code flowing "out like water through a burst dam," impossible for any human to fully understand [4]. Another, on the same setup: "I also fire off tons of parallel agents, and review is hands down the biggest bottleneck" [5]. Robert Laszczak put the asymmetry plainly: reviewing code written by an agent is much harder than reviewing code you wrote by hand [6]. None of the context is already in your head.
When the queue backs up, the review degrades into something else. Shmulik Cohen calls it vibe merging: developers overwhelmed by the volume skim the diff or hit Approve on gut feeling [7]. The tell is the LGTM speedrun, a 300-plus-line diff approved in under three minutes [8], which is a reading rate north of 100 lines per minute [18]. The skim itself is not new, and it used to be cheap: a fifty-line changeset you skimmed, you mostly still caught [16]. The skim has not changed. The changeset has.
The telemetry says the same thing. Faros AI read two years of data from 22,000 developers and found median time in code review up 441.5 percent, with pull requests merged with no review at all up 31.3 percent [9]. That first figure means the median PR now sits in review roughly 5.4 times as long as it did [17]. LinearB, across 8.1 million pull requests, found AI-assisted PRs run 2.6 times larger than hand-written ones and wait more than five times longer before a reviewer even picks them up [10]. Code arrives faster; reading does not scale to match. And the defect profile is the worst possible fit for a fast skim: in Stack Overflow's 2025 survey, 46 percent of developers distrust the accuracy of AI output against 33 percent who trust it, only 3 percent highly, and the top frustration, cited by 66 percent, is the answer that is "almost right, but not quite" [11].
The tempting fix is to automate the reading, and Opus 5 will volunteer for it: ask how to keep its mistakes out of main and it offers to review its own diffs [14]. For the mechanical layer, the AI reviewers now shipping on every code host do help, in the way a linter helps, by clearing the local feedback that would otherwise bog a human down before they reach the big-picture items [12]. Load-bearing approval is a different ask. A model's verdict on a model's diff is the same kind of stochastic guess that produced the code, on a question that gets harder with volume: per-line accuracy that looks fine on one function multiplies down across a hundred of them [13]. This is the same failure mode as a model grading its own instruction rules, where a stochastic judge run many times over gets less reliable, not more [15].
Worth watching: whether the merged-without-review rate keeps climbing [9], and whether anyone reports pickup latency falling rather than PR size rising [10]. Both are measurable in your own host. A useful local check is your median approved diff size against your median time-to-approve [8].