Build1 distinct publisher3 min readUpdated
A 46 percent merge rate is the first concrete number for handing routine upkeep to an agent. It is better than skeptics assume and nowhere near unattended.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
Anthropic has spent the last few weeks letting Claude Code perform daily maintenance on its own in-house apps, and there is now a number attached to the result: 388 pull requests opened, 180 merged [1][2][3][5]. That works out to a merge rate of about 46 percent [4], which is the first concrete figure operators have for what happens when routine upkeep is handed to an agent.
The number cuts both ways, and the arithmetic is the interesting part. If 180 of 388 landed, 208 did not [16], roughly 1.2 discarded pull requests for every one that shipped [18]. Every one of those still consumed a review slot: the merge decision came after a combination of automated Claude Code review and human review, according to Boris Cherny, the Anthropic engineer who created Claude Code [4][5]. The cost of this arrangement is not inference, it is reviewer attention, and at a 46 percent hit rate the agent generates slightly more review work that ends in nothing than review work that ends in a merge. Anthropic appears to know it: Cherny says the team is looking at ways to speed up the merge process for mechanical changes [14].
The setup is less exotic than the output suggests. Claude runs through a dedicated Slack channel called "proj-claude-maintains-apps," via a tool Cherny refers to as Tag, across iOS, Android, desktop, web, CLI, and the Agent SDK [7]. Twelve routines cover the upkeep surface [8]. A Crash Fuzzer opens apps in a simulator, taps around at random to provoke a crash, analyses the root cause and writes a fix [9]. A Dup Unifier looks for near-identical abstractions and proposes merging them [10]. A Dead-Code Remover deletes statically unreachable code, and for code it is only suspicious about, it first adds logging and checks the next day whether the path is actually cold [11]. That last one is the detail worth copying: the routine is designed to defer its own risky decisions by a day rather than guess.
Cherny also says there is no elaborate prompt engineering behind any of it, and shared prompts written in plain language instructing Claude to start daily crash-fuzzing routines on iOS, Android and desktop, use real apps rather than mocks, trigger crashes, and open pull requests with fixes [12]. The tuning loop lives at the routine level, not the prompt level: Claude usually gets a pull request right first time, and when it does not, the team adjusts the routine so the next day goes better, which sometimes takes a few days [13]. Cherny calls the results "surprisingly positive" and the experiment "early signs of life" for autonomous app maintenance [6][15].
Treat all of this as single-sourced vendor self-reporting on a vendor product. Worth watching: the denominator behind "a few weeks," since the source gives no exact duration and therefore no daily pull request volume [5]; whether the 46 percent holds as routines move past mechanical deletions and deduplication; the revert rate on the 180 that merged, which nobody has published; and what the faster merge path for mechanical changes actually removes from review [14].
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Anthropic is testing whether Claude Code can handle daily maintenance of the company's own software.
Claude created 388 pull requests across Anthropic's repositories in the first few weeks, according to Boris Cherny.
Of those pull requests, 180 were merged after human review.
After a combination of automated Claude Code review and human review, 180 pull requests were merged, a rate of about 46 percent.
Claude has been running daily maintenance on Anthropic's in-house apps for "the last few weeks," according to Boris Cherny, the Anthropic engineer who created Claude Code; no more precise duration is given.
Cherny calls the results "surprisingly positive."
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Specific numbers, single first-party channel
The core figures are unusually precise for agent-autonomy reporting (388 opened, 180 merged, twelve routines, named routine behaviours) and the outlet quotes the prompts and setup directly. But all of it traces to one internal Slack post by the creator of Claude Code, relayed by one publisher, with no repository links, no per-PR detail, no human baseline and no independent confirmation. Nine of the twelve routines are not described at all.
First-party internal pilot only
Adoption is real but confined to the vendor's own repositories: one company, a few weeks, twelve routines, six platform surfaces, with every change still passing through human review. No external customer, third-party team, or open-source project deployment appears in the supplied material, and Cherny himself frames it as an experiment showing 'early signs of life'.
Slightly overstated, but self-limited
Mild positive gap. The framing of an agent 'running daily maintenance' on production apps outruns what the numbers show — a supervised PR firehose where more than half the output is discarded and Anthropic is still hunting for ways to merge mechanical changes faster. The gap stays small because both the source engineer and the publisher label it an early experiment and report the rejection share rather than only the merge rate.
Vendor demonstrating its own product
The disclosure comes from the Anthropic engineer who created Claude Code, describing Claude Code succeeding on Anthropic's own codebases, published through his own Slack post and then relayed by a trade outlet that solicits subscriptions on the page. Every number in the cluster is chosen and released by the party that benefits commercially from Claude Code looking capable, with no adversarial or third-party check.
Facts clear, significance unresolved
Confidence in what was said is high — the figures and setup are stated plainly and are arithmetically consistent. Confidence in what it means is low: one publisher, one self-interested source, a few weeks of data, no baseline, no regression or reviewer-cost data, and no evidence the pattern transfers outside Anthropic's repos.
build
The $559M-versus-$12.3B quarter matters more than the $65B run rate4 distinct publishers
leadership
Slack Code makes the chat window a coding surface, and a platform call for engineering leaders1 distinct publisher
product
Slack Code makes coding agents taggable teammates. Your merge-approval policy is now overdue4 distinct publishers
build
Anthropic's CCAR-F puts a scaled score on "can build agents"1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 14, 2026