Build1 distinct publisher3 min readPublished
The aggregate mean belongs to a small cohort whose pipelines average six seconds. Read the median instead and a different picture appears, one where the trunk fails almost a third of its runs and needs over an hour to recover.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Follow any of these and your For You feed starts watching them — no settings page required.
security
The 2,500-org compromise was a Trivy problem. LiteLLM was the closing act.1 distinct publisher
invest
Cursor ships Origin to paying users as GitHub's outage count reaches 2571 distinct publisher
invest
OpenAI's own timeline: twelve days from agent attack to knowing it was them1 distinct publisher
build
A file-copy Allure adapter for Katalon, and the history IDs that make retries useful1 distinct publisher
A workflow run is one pipeline execution, and the counter increments by the same amount for a merge, a rerun, a scheduled job or a retry after a flake [16][11]. That is what breaks the productivity reading of the divergence. Feature-branch throughput up 15% against main-branch throughput down 7% is a 22-point spread [3][6], and it fits both a world where more work reaches the trunk and a world where each change that reaches it takes more attempts. The dev.to analysis says so plainly: if AI-assisted development produces more churn per shipped change, the counter climbs while delivery stands still, and the published data cannot separate the two cases [17].
Now the top of the distribution, because that is where the mean comes from. Fewer than one team in twenty grew code creation and code delivery together [6]. The top decile grew main-branch throughput by 1% [7]. Only in the top 5% does main-branch throughput grow 26%, alongside an 85% rise in feature-branch activity and daily runs going from 6.8 to 13.4 [8]. Rob Bowley's critique goes after the denominator: that cohort averages a CI pipeline duration of six seconds [14]. Multiply it out and the elite cohort executes roughly 80 seconds of pipeline per day, 13.4 runs at six seconds each [3]. For their doubling to say anything about your delivery, your gate would have to be doing what theirs does, which the source characterises as a lint job or a status check rather than a test suite [15]. If your main-branch gate is a fifteen minute integration suite, the number does not transfer, and the arithmetic is why.
The stability figures deserve the same discipline. A 70.8% main-branch success rate means 29.2% of main-branch runs ended red [9][1], which sits 19.2 points under CircleCI's own recommended 90% [9][2]. Reruns, scheduled jobs and retries all count in that population, and a red can be a dead runner, so the figure describes run outcomes rather than three in ten merges breaking the trunk [11]. Median time back to green of 72 minutes, up 13% year over year, implies about 64 minutes a year earlier [10][5]. It measures recovery time alone; nobody measured how long anyone actually sat blocked waiting [12].
Credit where the method earns it. Filtering to projects with at least two contributors and workflows that ran at least five times strips out toy repos and one-shot pipelines before anything gets averaged [5]. That filtering step is what makes the gap a causal one rather than a statistical artifact. Thoughtworks' complaint is that throughput without stability is waste rather than productivity, and that the report stops short of asking why builds fail more often [13]. It is also single-vendor telemetry, and the published methodology does not establish that the teams measured in September 2025 are the teams measured a year earlier [18][4].
The comparison figure I would carry into a review is the median gain of 4%; the headline mean runs close to fifteen times that figure [4]. In my context the number worth instrumenting is attempts per merged change, which no workflow counter can produce, because the counter cannot tell a passing merge from the third try at one. Everything else in the report is a well-filtered census of pipeline executions [5][16], and a trunk that fails 29.2% of its runs is a validation problem that needs no story about agents to explain it [1][9].
Ranked by verification strength, evidence, and original report placement.
The 59% figure measures the year-over-year increase in the average number of daily workflow runs across all CircleCI projects, on every branch; it is an aggregate average rather than anything branch-specific.
CircleCI's own report puts the median team's total throughput increase at 4%, and the bottom quartile saw no measurable increase at all.
For the median team, feature-branch throughput rose 15% year over year while main-branch throughput fell 7%.
CircleCI published the report on February 18, 2026, built from 28,738,317 workflows run during September 2025.
The dataset was filtered to projects with at least two contributors and workflows that ran at least five times, which removes toy repos and one-shot pipelines.
Fewer than 1 in 20 teams scaled code creation and code delivery at the same time.
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 28, 2026
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One relayed telling of one vendor's dataset
Every figure that matters — 28,738,317 workflows, 70.8%, 72 minutes, 6.8 to 13.4 daily runs — reaches a reader through a single dev.to post summarising a CircleCI report that nobody else in our coverage has opened. The post is unusually candid about its own limits: it publishes the filters, separates runs from merges, and says outright that it cannot size cohort churn. Candid is not corroborated, and the six-second pipeline duration doing the heaviest analytical work is second-hand from Rob Bowley with no primary version and no vendor reply.
Deployment-scale data about an undefined crowd
The measurement has real reach: 28.7 million workflows from projects with at least two contributors is not a survey panel, and the stability decline is drawn from the same population rather than asserted. What it lacks is a stable denominator of teams. September 2025's CircleCI users are compared against a year-earlier group nobody has shown to be the same teams, and everyone who moved to GitHub Actions, Buildkite or something in-house left the sample quietly. Wide, and about people we cannot name.
The travelling number dwarfs the typical one
The overstatement is measurable inside one report: 59% is what circulated, 4% is what the median team saw, and the cohort responsible for the difference finishes its pipelines in six seconds. Meanwhile the finding nobody amplified — a trunk failing almost three runs in ten and needing over an hour to recover — requires no AI narrative at all. dev.to is the correction rather than the hype, so the gap scored here belongs to the framing it dismantles; the writing in front of the reader is, if anything, more hedged than its own evidence demands.
Telemetry that concludes you need more CI
CircleCI produced the data, and CircleCI's CTO Rob Zuber reads it as proof that teams with autonomous validation are running laps around those who cannot validate AI-generated code at scale — a coherent story that also sells continuous integration. The report's silence on why builds fail sits exactly where a vendor's interest would place it. The counterweights in this reporting, a consultancy and an independent critic, are relayed rather than commissioned and have nothing riding on the conclusion.
Trust the direction, not the decimals
Two forces pull against each other. The shape — more branch activity, less merging, a shakier trunk — comes with denominators and caveats a reader can argue with, including the warning that red runs are not broken merges and that recovery time is not blocked time. But it is one writer reading one vendor's numbers, the six-second figure is borrowed, and the AI causation the report implies has no segmentation behind it. Quote the trend to a colleague; do not put 70.8% in a board deck yet.