Build1 distinct publisher3 min readPublished
The public software-factory writeups agree about the pipeline and disagree about who merges. That disagreement is the staffing question, because agents made code cheap to generate and left the cost of reading it alone.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Divide Stripe's throughput by a work week and the staffing question does most of the work for you. More than 1,000 PRs a week [1] is about 200 per working day [1]. Give each one ten minutes of human attention and that is roughly 33 hours a day, four people doing nothing else [2]. Ten minutes is generous for a dependency bump and absurd for a schema migration, which is the problem: the factory decides what arrives, and the review queue does not.
The pipeline is not exotic. An event arrives from an issue tracker, an error report or a Slack mention, one agent triages it, one reproduces and implements in a sandbox and runs tests, one reads the diff and scores risk, and a PR comes out the other end [15]. Vercel's version runs one agent each for classifying, analyzing, implementing and reviewing, closes the majority of incoming issues on its own, and keeps a human on every merge [8]. That last clause is the throughput ceiling, and it was installed on purpose. By mid-2026 the AI SDK repo was taking 100-plus new issues a month against over 1,000 open issues and nearly 800 open PRs [7], which is roughly ten months of intake already in the queue [7].
The best engineering in the set is Cloudflare's triage pipeline for Astro. It reproduces the bug, ships a preview build, waits for the original reporter to confirm the fix, and only then opens a PR [9]. The verification oracle is the person who filed the issue, who already has the repro and a reason to care. Astro's open issue count went from 200-plus to around 30 [10], an 85% cut [9]. PostHog answers the same constraint from the other end: StampHog auto-approves about a fifth of PRs for roughly $300 a month in tokens [11]. A fifth is a sampling rate, and somebody had to decide which fifth.
Uber's figures are the ones worth auditing, because they are the only ones about money. Weekly agentic requests grew 9.4x between February and August while total AI spend stayed roughly flat, which Uber credits to optimizing cost per session rather than letting usage run [4][6]. Flat spend across 9.4x volume implies cost per request fell about 89% [4]. It arrives with 3,600-plus reusable agent skills firing 30,000-plus times a day [5], an average of eight firings per skill per day [6], so most of them are narrow by construction. For that 89% to transfer, you need someone whose job is cost per session and a skills library specific enough to keep each context small.
The share numbers are drifting up quickly. Shopify's River co-authors one in eight merged PRs company-wide, 12.5% [2][3]. Ramp's Inspect is reportedly around 75% of merged PRs, up from 30% in January [12], 45 points inside a year [8]. Uber attributed more than 70% of PRs to local or cloud agents in its late-August 2026 post and said the number was still climbing [3]. Warp sells the pipeline as config, Terraform-for-agents, and claims about 30% of its own internal tasks run through it [13]. Every one of these has a name, from Minions to StampHog [16], which appears to be step zero.
BCG Platinion's 3-5x productivity numbers are self-reported by the organizations claiming them [14]. Read them as claims about someone else's workload until somebody publishes revert and escape rates beside merged-PR share. In my context I would copy Astro's ordering first: no PR exists until a human with the repro has confirmed the fix. The rest of the list mainly adds artifacts to the review tab.
Ranked by verification strength, evidence, and original report placement.
Stripe ships over 1,000 pull requests a week through an internal system it calls "Minions".
Shopify has a Slack bot named River that co-authors one in eight merged PRs across the whole company.
As of Uber's late-August 2026 engineering post, more than 70% of its pull requests are attributed to local or cloud agents, and the number is still climbing.
Uber's weekly agentic requests grew 9.4x between February and August.
Uber engineers have built over 3,600 reusable "agent skills" that fire more than 30,000 times a day, covering code review, CI self-healing, on-call triage and bug debugging.
Uber's total AI spend has stayed roughly flat over the February-to-August stretch, because it has been aggressively optimizing cost per session rather than letting usage run wild.
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · September 1, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
leadership
At Ramp, 75% of merged PRs come from a harness no vendor sold it2 distinct publishers
build
AI built the store in weeks. Production itemised what the apprenticeship would have cost1 distinct publisher
build
Flue 2 bets that agents are a rendering problem, not an orchestration one1 distinct publisher
product
A 2x LLM bill is not a bug report: token spend is an observability problem1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One byline behind every number
Seventeen quantified claims and a single author standing behind all of them. What saves it from being weaker: most figures point at artifacts a reader can go inspect — Uber's engineering post, Vercel's public factory repo, Astro's issue tracker, triagebot-action on GitHub. What holds it down: no second newsroom has touched any of it, and dev.to hedges its two softest numbers itself, calling Ramp's 75% "reportedly" and BCG's 3–5x self-reported. Definitions are the real weak point — "attributed to agents" and "co-authored" are counted by whoever benefits from the count.
Ten companies deep, no denominators
This is the strongest leg of the story. Agent pipelines are not a pilot here: Uber past 70% of PRs, Shopify at one merge in eight, Stripe over a thousand PRs a week, Astro's backlog down to about 30 open issues, and two pipelines released for anyone to run. Warp is selling the pattern as a product. The ceiling on the score is that every share is a numerator without a stated denominator — no company says how many PRs total, how large the change, or what an agent had to contribute to earn attribution.
Mildly overstated, and it says so itself
The word "factory" does more work than the evidence licenses — a PR count is throughput, not shipped value, and nobody in this reporting measures what the merged changes were worth. But the overstatement is modest because dev.to keeps puncturing itself: it labels the consultancy multiples self-reported, hands the microphone to a four-month lights-off run that left a codebase rotted enough for one bug to take weeks, and points at GitClear's independent line-level analysis showing rising duplication. The residual gap is the unexamined leap from "agents wrote the PR" to "engineering got 3–5x faster".
Recruiting posts, vendor decks, one honest cost line
Trace each figure to who published it. Uber, Stripe, Shopify and Ramp are employers advertising engineering capability. Vercel, Warp and Factory.ai sell into the category they are describing. BCG Platinion is naming an era it can bill against. That is nearly the whole record. The exceptions are the numbers that cost their authors something to admit: PostHog's $300-a-month token bill and Uber's flat spend against 9.4x volume are cost disclosures, which are harder to dress up than throughput percentages. The relaying author has no visible commercial stake, which is why the caveats survived into print.
Coherent, single-sourced, and vague where it counts
The internal logic holds up: cheap generation, expensive reading, a pipeline to absorb the difference, and a real fault line over who merges. It is also one account, relaying primary posts nobody in our coverage has independently checked, and it goes quiet exactly where the story's own thesis lands — the reviewer side. If the bottleneck is now whoever reads the diff, the number that matters is reviewers and minutes per PR, and no company here publishes either.