BuildNot yet confirmed elsewhere1 publisher2 min readPublished
Texas Tech study finds coding-agent revert rates range from 6.1% to 14.5% by vendor
Obada Kraishan of Texas Tech tracked 37,623 GitHub pull requests and found 90-day revert rates of 6.1% for OpenAI Codex and 14.5% for Devin. A pooled 'AI code' rate from that sample mostly measures Codex, so each tool needs its own numbers.
The Engineer · Build desk
What happened
- Human-written PRs drawn from the same repositories over the same window had 11.5% of their changes reverted within 90 days of merge.
- Codex and Devin are the only vendors whose confidence intervals clear the human rate; Copilot at 12.5%, Cursor at 11.4% and Claude Code at 10.5% show no detectable difference.
- Pooled agent code was less likely than human code to contain a security smell, at an odds ratio of 0.63, but the whole effect sat in the largest PRs.
- Claude Code PRs waited longest for a first human review, a median 12.6 hours, and were the largest in the sample at a median 495 changed lines.
Why it matters
- decision A review policy set at the pooled agent rate would relax scrutiny for every vendor except the one that produced most of the sample.
- constraint Copilot, Cursor and Claude Code cannot be ranked on revert risk from this data, so choosing among them needs a team's own labelled PRs.
- constraint The security advantage holds only for very large diffs, so it gives no grounds for lighter security review of typical-sized agent PRs.
- cost Teams that let agents open very large PRs pay in review latency, and they can measure that cost from their own queue data before any vendor comparison.
The odds ratios in the write-up are the raw revert rates restated. Convert Codex's 6.1% and the human rate to odds and the ratio comes out at 0.50. Devin's 14.5% gives 1.31 [27]. Both match the listed figures, and the other three rows match to within rounding [27]. The study is observational and nothing was randomized, so each vendor's row measures the tool together with the work its users chose to give it [4]. As reported, the revert comparison looks unadjusted apart from one control: the human PRs come from the same repositories over the same window [2][27].
I weighted each vendor's rate by the sample size listed beside it. That gives a pooled agent revert rate of about 7.7%, well under the humans [24]. Codex supplies 17,756 of the 33,596 agent PRs, about 53% [12][23]. Take Codex out and the other four vendors pool to about 13.0%, above the human rate [25]. Whether pooled agent code reverts more often than human code depends on whether one product is in the sample [24][25]. Among the agents, Codex and Devin sit 8.4 percentage points apart. Codex and the humans are 5.4 points apart [26]. The write-up's author wrote that "the pooled agent baseline is mostly Codex behavior wearing a trench coat" [13].
The rates transfer to a team only if its code resembles public GitHub repositories with more than 100 stars [4]. Session length matters too. The write-up argues that per-turn instruction compliance decays as an agent session runs longer, and that the paper does not control for session length. On that argument, a vendor whose users run longer sessions is penalized or flattered by a variable nobody measured [14]. Sample size sets a further limit. Claude Code's row rests on 267 PRs with an odds-ratio interval of 0.60 to 1.36, and Cursor's rests on 946 [10][9].
The part a team can copy is the instrumentation. The write-up recommends labelling each PR by author type (agent, human or mixed) when it is created, because retrofitting the label later is guesswork [19]. The paper's result depends on every PR keeping its vendor label [21]. I'd record the vendor in that label as well as the author type. I'd also rank tools only where a team's own interval clears its human baseline. The author wrote: "You do not need 2,807 repositories. You need provenance and a fixed window." [20]
What to watch
- Whether the published version of the paper adds revert models adjusted for PR size and task; a Devin gap that survives those adjustments would point at the tool and not its users' workload.
- Replication of the per-vendor revert rates on private or internal codebases outside the public-repo filter.
- Larger samples for Claude Code and Cursor that narrow their intervals enough to rank them against humans.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+10
- Incentives35
- Confidence50
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
A preprint by Obada Kraishan at Texas Tech, 'Not All Agents Are Equal: Code Quality and Post-Merge Maintenance Across Five Autonomous Coding Agents in the Wild', follows 37,623 provenance-labeled pull requests across 2,807 GitHub repositories from December 2024 to July 2025.
- [2]
Of the PRs, 33,596 were opened by five commercial agents (OpenAI Codex, Devin, GitHub Copilot, Cursor, Claude Code) and 4,027 by a matched human baseline in the same repositories over the same window.
- [3]
The study follows each merged change for 90 days; post-merge maintenance covers 26,283 merged PRs, measuring size-normalized churn and revert detection.
- [4]
The design is observational, nothing was randomized, and the repositories are public repos with more than 100 stars.
- [5]
Within 90 days of merge, the human baseline reverted 11.5% of its PRs.
- [6]
OpenAI Codex PRs had a 6.1% revert rate within 90 days (n = 17,756, odds ratio 0.50, 95% CI 0.44 to 0.57).
- [7]
Devin PRs had a 14.5% revert rate within 90 days (n = 2,185, OR 1.31, CI 1.11 to 1.54).
- [8]
GitHub Copilot PRs had a 12.5% revert rate (n = 2,094, OR 1.10, CI 0.93 to 1.31, p = .457).
- [9]
Cursor PRs had an 11.4% revert rate (n = 946, OR 1.00, CI 0.79 to 1.25).
- [10]
Claude Code PRs had a 10.5% revert rate (n = 267, OR 0.90, CI 0.60 to 1.36).
- [11]
Devin and Codex are the two rows whose confidence intervals clear the human baseline; Claude Code, Copilot and Cursor show no detectable difference from humans.
- [12]
Codex accounts for 17,756 of the 33,596 agent PRs.
- [13]
"the pooled agent baseline is mostly Codex behavior wearing a trench coat"
- [14]
Per-turn instruction compliance decays the longer an agent session runs, and session length is not controlled in the study, so a vendor whose users run longer sessions gets penalized or flattered by a variable nobody measured.
- [15]
Pooled agent code was less likely than human code to contain a security smell (OR 0.63), driven by fewer hardcoded credentials and eval-style constructs; a size-stratified check puts the whole effect in the largest PRs (XL bucket, delta -0.08, p = .025) with no difference in the four smaller buckets.
- [16]
Patch-level quality analysis covers 8,933 PRs and 1,348,822 added lines, scoring security smells across eight CWE classes for Python, JavaScript and TypeScript.
- [17]
Claude Code PRs waited the longest for a first human review, median 12.6 hours, and were the biggest: median 495 changed lines, median nesting four levels deep, highest branch density at .063 per line.
- [18]
Copilot PRs drew the most human reviews and change requests.
- [19]
The write-up recommends labelling PRs by author type (agent, human, mixed) at creation time, saying retrofitting this later is guesswork.
- [20]
"You do not need 2,807 repositories. You need provenance and a fixed window."
- [21]
Every PR keeps its vendor label, so the analysis can separate agents from individual products, and that distinction carries most of the result.
- [22]
Devin's 90-day revert rate is about 2.4 times Codex's.
- [23]
Codex supplies about 53% of the agent PRs.
- [24]
Weighting each vendor's revert rate by its listed n gives a pooled agent revert rate of about 7.7%, below the human 11.5%.
- [25]
Excluding Codex, the other four vendors pool to a revert rate of about 13.0%, above the human 11.5%.
- [26]
The revert-rate gap between Codex and Devin is 8.4 percentage points, against 5.4 points between Codex and the human baseline.
- [27]
Odds ratios computed from the raw revert rates reproduce the listed ones: Codex 0.50 and Devin 1.31 exactly, Copilot 1.10, Claude Code 0.90 and Cursor 0.99 against a listed 1.00, consistent with an unadjusted comparison.
Sources
1 independent publisher whose own reporting we read for this story.
- dev.toAgent PRs revert at 6.1% or 14.5%, depending on the vendor
1 article · October 9, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.
Topics
- AI Coding AgentsFollow
- Empirical Software EngineeringFollow
- Code reviewFollow