Skip to content

BuildNot yet confirmed elsewhere1 publisher2 min readPublished

Texas Tech study finds coding-agent revert rates range from 6.1% to 14.5% by vendor

Obada Kraishan of Texas Tech tracked 37,623 GitHub pull requests and found 90-day revert rates of 6.1% for OpenAI Codex and 14.5% for Devin. A pooled 'AI code' rate from that sample mostly measures Codex, so each tool needs its own numbers.

The Engineer · Build desk

How we use AISend a correction

What happened

  • Human-written PRs drawn from the same repositories over the same window had 11.5% of their changes reverted within 90 days of merge.
  • Codex and Devin are the only vendors whose confidence intervals clear the human rate; Copilot at 12.5%, Cursor at 11.4% and Claude Code at 10.5% show no detectable difference.
  • Pooled agent code was less likely than human code to contain a security smell, at an odds ratio of 0.63, but the whole effect sat in the largest PRs.
  • Claude Code PRs waited longest for a first human review, a median 12.6 hours, and were the largest in the sample at a median 495 changed lines.

Why it matters

  • decision A review policy set at the pooled agent rate would relax scrutiny for every vendor except the one that produced most of the sample.
  • constraint Copilot, Cursor and Claude Code cannot be ranked on revert risk from this data, so choosing among them needs a team's own labelled PRs.
  • constraint The security advantage holds only for very large diffs, so it gives no grounds for lighter security review of typical-sized agent PRs.
  • cost Teams that let agents open very large PRs pay in review latency, and they can measure that cost from their own queue data before any vendor comparison.

The odds ratios in the write-up are the raw revert rates restated. Convert Codex's 6.1% and the human rate to odds and the ratio comes out at 0.50. Devin's 14.5% gives 1.31 [27]. Both match the listed figures, and the other three rows match to within rounding [27]. The study is observational and nothing was randomized, so each vendor's row measures the tool together with the work its users chose to give it [4]. As reported, the revert comparison looks unadjusted apart from one control: the human PRs come from the same repositories over the same window [2][27].

I weighted each vendor's rate by the sample size listed beside it. That gives a pooled agent revert rate of about 7.7%, well under the humans [24]. Codex supplies 17,756 of the 33,596 agent PRs, about 53% [12][23]. Take Codex out and the other four vendors pool to about 13.0%, above the human rate [25]. Whether pooled agent code reverts more often than human code depends on whether one product is in the sample [24][25]. Among the agents, Codex and Devin sit 8.4 percentage points apart. Codex and the humans are 5.4 points apart [26]. The write-up's author wrote that "the pooled agent baseline is mostly Codex behavior wearing a trench coat" [13].

The rates transfer to a team only if its code resembles public GitHub repositories with more than 100 stars [4]. Session length matters too. The write-up argues that per-turn instruction compliance decays as an agent session runs longer, and that the paper does not control for session length. On that argument, a vendor whose users run longer sessions is penalized or flattered by a variable nobody measured [14]. Sample size sets a further limit. Claude Code's row rests on 267 PRs with an odds-ratio interval of 0.60 to 1.36, and Cursor's rests on 946 [10][9].

The part a team can copy is the instrumentation. The write-up recommends labelling each PR by author type (agent, human or mixed) when it is created, because retrofitting the label later is guesswork [19]. The paper's result depends on every PR keeping its vendor label [21]. I'd record the vendor in that label as well as the author type. I'd also rank tools only where a team's own interval clears its human baseline. The author wrote: "You do not need 2,807 repositories. You need provenance and a fixed window." [20]

What to watch

  • Whether the published version of the paper adds revert models adjusted for PR size and task; a Devin gap that survives those adjustments would point at the tool and not its users' workload.
  • Replication of the per-vendor revert rates on private or internal codebases outside the public-repo filter.
  • Larger samples for Claude Code and Cursor that narrow their intervals enough to rank them against humans.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence55
Adoption
Insufficient
Hype gap+10
Incentives35
Confidence50
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    A preprint by Obada Kraishan at Texas Tech, 'Not All Agents Are Equal: Code Quality and Post-Merge Maintenance Across Five Autonomous Coding Agents in the Wild', follows 37,623 provenance-labeled pull requests across 2,807 GitHub repositories from December 2024 to July 2025.

    ReportedSupportedSource: dev.to write-up of the preprintView cited source
  2. [2]

    Of the PRs, 33,596 were opened by five commercial agents (OpenAI Codex, Devin, GitHub Copilot, Cursor, Claude Code) and 4,027 by a matched human baseline in the same repositories over the same window.

    ReportedSupportedSource: dev.to write-up of the preprintView cited source
  3. [3]

    The study follows each merged change for 90 days; post-merge maintenance covers 26,283 merged PRs, measuring size-normalized churn and revert detection.

    ReportedSupportedSource: dev.to write-up of the preprintView cited source

Sources

1 independent publisher whose own reporting we read for this story.

  1. dev.to

    1 article · October 9, 2026

    Agent PRs revert at 6.1% or 14.5%, depending on the vendor

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Topics

Loading related stories