Build1 distinct publisher3 min readPublished
A 90-day instrumented study reports 28% more pull requests and 13 points of test coverage, but the gain lands only after training and workflow embedding, which puts the money on enablement rather than seats.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
Follow any of these and your For You feed starts watching them — no settings page required.
build
AI's 4x code generation ships with a doubled review cycle and tripled post-merge fixes1 distinct publisher
build
AI-written code fails the same four ways, and every gate you own reports green1 distinct publisher
product
Engineering counts merged pull requests and nothing for the hours spent watching the agent1 distinct publisher
build
Claude Code now outruns Copilot roughly two to one in JetBrains' survey of 15,000 developers1 distinct publisher
Weeks one and two show the clearest signal in the design. Access went out with no guidance at all, deliberately, so the team could watch organic adoption [5]. What the instrumentation caught was divergence. The author reports the prompting skill gap widened, and that the widening was largest in the four weeks before training ran [19].
The arithmetic behind that spread is stark. A senior with strong prompting at roughly 40% velocity improvement, a junior with weak prompting at roughly 5%, is an eight-to-one difference in outcome from the same tooling on the same codebase [19][2]. The author is explicit that the results came out more nuanced than the vendor benchmarks suggest [22].
So price the enablement. Two 90-minute sessions per squad is three hours of contact time per engineer, which across 70 engineers is about 210 engineer-hours if everyone attends both [6][1]. Add the weekly retros running through months two and three [7]. Add a second round of teaching, because in month one review times fell for the wrong reason: reviewers were approving faster rather than reading better, and correcting that took explicit re-training [18]. That cost lives outside the seat invoice.
For the 28% to transfer, several things have to be true of your team [9]. You need your own six-month baseline out of GitHub and Jira, or you are comparing against a memory [4]. You need the seniority mix that produced the gain, since engineers with two to five years of experience saw the largest boilerplate wins while seniors who had already internalised those patterns saw less [13]. You need coverage headroom, because the biggest reported win was test writing and coverage moved 13 points [11][12]. And you need the review culture rebuilt, which is the author's own headline finding: obvious errors fall, while plausible-but-wrong logic errors hold flat or rise if review does not adapt [17].
The coverage-versus-defects gap is the number I would interrogate first. Thirteen points of new tests bought a 7% improvement in escape rate [11][10], and two logged incidents ran the other way when engineers over-trusted generated domain calculations that were syntactically fine and semantically wrong [16]. Coverage is a lines-executed metric, a separate question from whether invariants are asserted. That reading is mine, not the study's.
One more limit on what this data can settle. Claude Code was primary, Copilot ran on a single squad, and Cursor on three frontend engineers, about 4% of the team [3][4]. There is no control group held still, participation was voluntary, and the two engineers who opted out had both opted back in by week six, which is the closest thing to a unanimous verdict in the dataset [8].
In my context this reorders the plan more than the budget: book the training dates first, then switch on access, because the pre-training window is where the reported damage shows up [19].
Ranked by verification strength, evidence, and original report placement.
A structured 90-day study of AI coding assistants was run across a 70-engineer enterprise team covering frontend, backend, platform and QA.
Metrics tracked were PR velocity, review cycles, defect escape rate, time-to-first-review and self-reported time savings.
Tools used were Claude Code as primary, GitHub Copilot on one squad for comparison, and Cursor with three frontend engineers.
A six-month baseline was established before any tooling changes, using PR metrics from GitHub, defect data from Jira and monthly velocity surveys.
In weeks 1 and 2 tool access was granted with zero guidance, in order to observe organic adoption patterns.
In weeks 3 and 4 the team ran prompt engineering training consisting of two 90-minute sessions per squad.
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 27, 2026
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Instrumented, but nobody has seen the export
A six-month pre-period built from GitHub and Jira puts this well above the usual AI-productivity anecdote, and the author is specific about which metrics moved. What is missing is everything that would let a reader push back: no control squad despite one squad running Copilot and three engineers running Cursor, no distributions or significance behind +28% and 13 points, and no examination of whether PR count still means the same thing once AI changes how work is chunked. The most-quoted number in the piece, the eight-to-one prompting spread, is introduced as what an engineer 'might see'.
One real deployment, one anonymous employer
This is not a pilot memo — 70 engineers used these tools in production work for a quarter, with voluntary uptake strong enough that both opt-outs reversed themselves by week 6. But it is exactly one org, unnamed, in an unstated industry, and the tool distribution is lopsided: Claude Code carried nearly everything while Cursor touched three people. Nothing in our coverage says whether this team's codebase, review norms or seniority mix resemble anyone else's.
Honest about defects, loose about return
Most of this post argues against its own headline: the defect improvement is called smaller than expected, two incidents are recorded where AI made things worse, and the falling review times of month 1 are named as a false win. That is unusually self-critical. The overshoot sits in the closing arithmetic, where self-reported hours multiplied by a blended $95 becomes a 32× return, with the enablement effort that the author credits for the gains left off the cost line. The 90%-haircut defence sounds rigorous but simply restates a number whose source is a survey.
Disclosure cut off mid-sentence
We cannot say who benefits from this account. The employer is never named, no vendor relationship or procurement stake is disclosed, and the text breaks off in the middle of what appears to be a disclosure line about AI-generated code. That is not evidence of a conflict; it is an absence of the facts needed to weigh one, so we are not assigning a number.
Credible mechanisms, unreplicated numbers
Our confidence splits along a clean line. The qualitative machinery — reviewers rubber-stamping AI output, plausible-but-wrong domain logic, context loading beating bare signatures — is coherent, specific and consistent with the failure modes practitioners describe elsewhere, and we would not be surprised to see it replicated. The percentages are another matter: one author, one team, one publication, no second reading anywhere in our coverage.