Build1 publisher3 min readPublished
Coding agents have made CI the bottleneck at Anthropic and Linear
Anthropic's CI job volume grew 25x in six months as coding agents raised pipeline load, The New Stack reports. Faster runners and test selection cut the cost of each run, yet repo tests still mock the service seams where distributed systems break.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Linear says its test suite has nearly quadrupled since January, that agents now write most of its tests, and that it reworked its pipeline end to end.
- DORA found that "higher AI adoption is associated with increases in both software delivery throughput and software delivery instability."
- Cursor says more than 30% of the PRs it merges come from agents running in cloud sandboxes, each with its own virtual machine.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- contradiction Anthropic's 25x works out to about 13% growth a week, above Blacksmith's 5-10% band, so the figure The New Stack calls typical sits above what a runner vendor sees across its fleet.
- constraint Test selection shrinks each run, but a repo's suite still mocks its downstream services, so a quicker green check leaves cross-service breakage as untested as before.
- decision Teams now have to choose between keeping verification as a gate after the PR and moving it into a sandbox the agent runs while it works, the route Cursor, Copilot and Codex have taken.
For about twenty years, CI was sized for human output, according to the New Stack analysis that collected these figures [7]. A developer opened a few pull requests a week. A 20-minute pipeline barely mattered, because the developer was already on the next task [7].
Agents break that sizing on volume first: one engineer running several agents in parallel raises PR count by a multiple [8]. They also break it on placement. CI runs after the PR exists. By the time a red check comes back 20 minutes later, the agent no longer has its working context, so every failure costs a full round trip [9].
The 25x is a fact about Anthropic's workload. The New Stack calls it typical and points to Blacksmith, which sells CI runners, reporting that the CI jobs it runs grow 5% to 10% week over week [10][6]. A runner vendor reporting rising runner demand is the least surprising data point in the piece. Compounded over 26 weeks, Blacksmith's band comes to roughly 3.6x to 11.9x [1][2]. Anthropic's 25x over the same span implies about 13% a week, above the top of that band [3]. The piece does not say whether Blacksmith's growth comes from existing customers or new ones. A team would need PR volume compounding near 13% a week to face Anthropic's load [3].
The 8x figure for code shipped per quarter uses a 2021-to-2025 baseline [2]. The 25x job figure uses a six-month window [1]. The two do not combine into a jobs-per-line ratio.
Test impact analysis is good engineering. Running only the tests a change could affect cuts the work per PR and goes straight at volume [3]. Linear reworked its pipeline end to end to keep up with the same kind of growth [4]. I think selection is the right first move where the repository is the whole system, as it is for a standalone application [18].
The New Stack author's objection is that a faster gate still checks the same thing, and what it checks is a repository [11]. In a cloud-native system, a repository is one service out of forty, and its tests mock everything else [12]. A change can pass every unit test, pass CI, clear a sandbox spun up from its branch, and still fail on the first live request that reaches into another service [12]. The cited cases include a response field renamed while a downstream consumer still reads it, and a timeout tightened in one service that cascades into retries in another [13]. DORA's finding fits the objection. Higher AI adoption is associated with increases in both delivery throughput and delivery instability [14].
Cursor moved verification into the agent's working loop. Each of its agents works inside a cloud sandbox on a virtual machine of its own, and agents working that way account for more than 30% of the PRs Cursor merges [15]. "Without the ability to use the software they are creating, agents hit a ceiling," Cursor wrote [16]. GitHub's Copilot cloud agent runs tests in an ephemeral environment powered by GitHub Actions, and Codex runs a setup script and resumes cached containers [17]. Depot's CEO wrote that the future is "giving agents a way to validate code and maintain trust as they work" [5]. A sandbox built from one branch is still the failure case the New Stack piece describes: the branch sandbox passes, and the cross-service request breaks [12].
What to watch
- Whether Anthropic or Linear publish follow-up data on queue time or escaped defects after test impact analysis and the pipeline rework.
- Whether DORA's next report puts a size on the delivery instability it associates with AI adoption.
- Whether agent sandboxes from Cursor, GitHub or OpenAI start provisioning the services a change calls, beyond the branch under test.