Skip to content

Build1 publisher3 min readPublished

AI-assisted pull requests wait 4.6 times longer for a first review than unassisted ones

LinearB's 2026 benchmark finds only 32.7% of AI-assisted pull requests accepted within 30 days, against 84.4% of manual ones. The figures are associations, but they put review cost in the queue and in repeat rounds, where faster diff reading helps little.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying AI-assisted pull requests wait 4.6 times longer for a first review than unassisted ones
Generated illustration

What happened

  • The same LinearB data, which CodeRabbit cites in its guide to reviewing AI-generated diffs, finds AI-assisted pull requests 2.6 times larger than unassisted ones.
  • Obada Kraishan's preprint covers 37,623 provenance-labeled pull requests opened by five autonomous coding agents across 2,807 repositories.
  • Claude Code pull requests in Kraishan's dataset waited a median of 12.6 hours for a first human review.
  • Copilot pull requests drew the most human reviews and the most change requests of the agents studied.
  • Codex pull requests were reverted 6.1% of the time against 11.5% for human ones, while Devin's were reverted 14.5% of the time.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision Choosing a review tool now turns on its trigger list: a review that fires on PR open overlaps the queue, while one that waits for green CI or a comment command adds its latency on top.
  • cost A tool that spreads ten findings across four review cycles charges the author four context reloads and four re-reviews for what one pass could have delivered.
  • exposure Teams picking an agent vendor also pick a revert rate, and on Kraishan's figures a Devin PR is about 2.4 times as likely to be reverted as a Codex PR.

Time-to-merge bundles at least three costs, and reading the diff is only one of them, according to the dev.to analysis that assembled these figures [10]. The other two are the wait before a first review and the rework after a change request [10]. When the number drops after a team adds an AI reviewer, the post argues, the drop usually comes from those two [10].

The Claude Code median is the strongest support for that claim. It comes to 756 minutes between a PR opening and a human looking at it [3]. The post's explanation is queue time: the agent opens the PR, and nothing happens until a person picks it up [14]. Reviewers do not read one diff for twelve hours [14]. A tool that reads faster than a human trims the few minutes at the end of that window and leaves the hours at the front where they were [14].

The acceptance gap is weaker evidence for the same claim. An acceptance rate of 32.7% leaves 67.3% of AI-assisted PRs unaccepted after 30 days [1]. The manual figure is 15.6%, so an AI-assisted PR is about 4.3 times as likely to be unaccepted at the 30-day mark [2]. The post takes that to mean the cycle spans several passes per change [17]. A PR closed after one round and a PR in its fourth round count the same in that figure. CodeRabbit, which cites the LinearB numbers, labels them observed associations and makes no causal claim [4].

None of the published figures compare a team before and after it adopted a review tool. LinearB compares AI-assisted PRs with unassisted ones, and Kraishan compares agents with each other and with human authors [1][9]. So the case that review-tool gains come from less waiting and less rework is an inference from where agent PRs spend their time. I think the inference is sound for the wait segment and still open for rework.

The post's remedy is to record four timestamps per PR [11]:

1. The PR opens. 2. The first review lands. 3. The first change request lands. 4. The PR is approved and merged.

Wait is the gap from 1 to 2, reading runs from 2 to 3, and rework is everything after 3, including the author's fix and each re-review [11]. Split by provenance and by repository, the data then picks the lever [16]. If wait dominates, change the trigger [16]. If rework dominates, judge whether the tool files one precise change request or a stream of comments the author answers over several passes [16]. Reading speed is the right purchase only when reading dominates and diffs are large [16]. I would run it in that order, with the vendor demo after the team's own data.

Kraishan's study is a preprint and has not been peer reviewed [6]. It did release its pipeline code for replication, which is the part of the work I would copy [5]. A team can run the provenance labelling against its own repositories and check whether its agents queue the way Claude Code's did [5][8].

The other lever sits before review starts. In a Stack Overflow post on coding guidelines for AI and people, Heroku chief architect Vish Abrams makes the point that principles seasoned engineers assume, such as DRY, are not common knowledge to an agent [15]. The dev.to post argues that written rules move the standards check earlier, so a reviewer's first pass is not spent on naming and layout, the work that generates change requests [18].

What to watch

  • Replication of Kraishan's preprint with its released pipeline, and whether the 12.6-hour median holds outside its 2,807 repositories.
  • Before-and-after time-to-merge data from LinearB or CodeRabbit for teams that adopted a review tool, split into wait, reading and rework.
  • First-review wait figures for the other four agents in Kraishan's dataset, to show whether Claude Code's queue is typical.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories