Build1 publisher3 min readPublished
The defect tax on in-editor models is a review capacity problem, not a tooling one
Reported figures put bug rates 41% higher and security flaws 2.74x more prevalent in AI assisted code. The constraint that decides whether that matters is review throughput.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- One study, attributed to Uplevel, found a 41% increase in bug rates among teams that adopted AI coding tools.
- The mechanism identified in the Uplevel finding was over-reliance: developers taking suggestions without enough review behind them.
- Static analysis across enterprise repositories found security vulnerabilities roughly 2.74x more prevalent in AI assisted code than in manually written code; the write-up does not name the study.
- Around 95% of code reaching production now contains AI assisted content.
- CodeRabbit's research found AI assisted codebases carried 1.7x more issues per pull request than human-only code, measured per pull request rather than per line.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
A dev.to write-up collecting the current numbers on AI assisted coding puts the post-adoption bug rate increase at 41%, citing Uplevel [1], and security vulnerabilities at roughly 2.74x the prevalence found in manually written code, citing static analysis across enterprise repositories [3]. The second figure is the one that changes planning, because security defects carry a cost curve that ordinary bugs do not [19].
The adoption argument is over. The same piece reports that around 95% of code reaching production now contains AI assisted content [4], which leaves roughly 5% that a review policy scoped to "the AI code" would actually cover [18]. Ask engineers whether the tools make them faster and the hands go up; ask whether the code is better and they drop [13]. Both of those are honest answers and they are not in conflict.
The mechanism Uplevel identified was over-reliance: suggestions taken without enough review behind them [2]. CodeRabbit's research approaches it from the other side and finds AI assisted codebases carrying 1.7x more issues per pull request than human-only code [5]. Per pull request is the load-bearing detail, because the PR is the unit your team actually reviews. That is about 70% more issues arriving at each review gate [16]. If the gate did not get 70% more attention, it is now catching a smaller fraction of what is there. Nothing mysterious is happening; the volume changed and the filter did not.
The underlying inversion is that writing code got cheaper than reading it, roughly at the same moment across Cursor, Copilot and Claude Code [7]. A tool that emits 300 lines in one shot does not emit 300 lines of review attention beside it, and nobody budgeted for the difference [8].
What lands is also unglamorous. Not exotic hallucination but an error path that swallows the exception, or a null check that reads correctly and sits one branch too late, in code that passes review because it looks like code you would have written [9]. On the security side, a model produces what is statistically typical for the surrounding context, and typical public code includes string concatenation into queries, permissive CORS, secrets read from a literal, and validation that trusts an object's shape because the type annotation said so [10]. The model also does not know where your trust boundary sits, or that a given handler is reachable without auth [11]. Type safety does not close that gap: type safe code can authorise the wrong user [12].
Comparing the two headline figures, security flaw prevalence is running at roughly twice the multiplier of the overall bug rate [17]. That is a mix shift, not just a volume shift.
The reason teams do not self-correct is timing. The speed is felt in the same session; the defect surfaces weeks later in an incident channel, usually attributed to something else, so the feedback that would calibrate trust never arrives [14]. Trust tends to break in one specific moment: when someone traces a production incident to a block of code nobody on the team can explain [15].
Two things to watch. First, provenance: the 41% and 2.74x figures travel widely as citations, and the write-up does not name the enterprise repository study behind the second one [3]. Second, your own issues-per-PR trend and the share of those issues that are security class, which is the only version of these numbers you can act on.