Skip to content

Build1 publisher3 min readPublished

Two AI coding experiments point opposite ways because they measured different populations

A field study of 4,867 developers recorded 26.08% more tasks completed with GitHub Copilot, while METR timed 16 experienced developers taking 19% longer in repositories they already knew. The two figures count different things.

The Engineer · Build desk

Photograph accompanying Two AI coding experiments point opposite ways because they measured different populations
Photo: metr.org

What happened

  • Field experiments at Microsoft, Accenture and an anonymous Fortune 100 company gave GitHub Copilot access to some of 4,867 developers and recorded 26.08% more tasks completed by those with access.
  • The same research found the larger productivity gains among the less experienced developers in the sample.
  • METR's July 2025 randomised trial timed 16 experienced open-source developers on 246 tasks in mature repositories they knew, and they took 19% longer when early-2025 AI tools were allowed.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint A count of completed tasks does not convert into hours without the task definition, so 26.08% cannot be entered in a capacity plan as a 26% reduction in delivery time.
  • decision Rolling out seats uniformly now argues against both measurements at once, because the gain was concentrated in less experienced developers and the slowdown was measured in experienced ones working in code they knew.
  • exposure Any programme whose success metric is developer sentiment is measuring the wrong variable: in METR's trial, belief and stopwatch disagreed by about 39 percentage points.
  • contradiction Quoting the 19% slowdown as a current figure puts a leader at odds with METR itself, which says its newer data cannot pin the present effect down.

The two headline numbers are not in the same units. The 26.08% is a count: developers given access to GitHub Copilot finished more tasks than developers without access, across three workplace experiments covering 4,867 people [1][2]. The 19% is a clock reading: 16 experienced open-source developers took longer per task when early-2025 AI tools were allowed than when they were not [4][5]. A task count and a completion time answer different questions, and neither study supplies the bridge between them.

For the 26.08% to land on your team, your task mix has to resemble the mix at the three sites, your work has to be countable in tasks the way the study counted it, and your population has to skew the way the gain did. The researchers found that less experienced developers gained more [3]. The post reporting the result does not describe how a completed task was defined [18]. It also quotes the figure to two decimal places. I would not plan against the second one.

For the 19% to land, the conditions are narrower and easier to check. METR's participants worked in mature repositories they already had experience with, on 246 tasks, about 15 each [4][14]. In that setting the cost of accepting AI output is highest, because the developer still has to understand the generated code, check whether it works, test it, find the mistakes, integrate it into the existing codebase and sometimes correct the suggestion [12].

The self-report figures are the part I would take into a planning meeting. Before the trial the developers expected AI to reduce their completion time by 24% [6]. After finishing the tasks they still believed it had made them about 20% faster [7]. The measurement was 19% slower [5]. That leaves 43 percentage points between forecast and outcome, and about 39 between post-task belief and measurement [15][16].

The dev.to author who assembled both results does not treat them as a conflict. "I don't think these findings necessarily contradict each other," the author wrote, and pointed to the different developers, environments, tasks and AI-tool contexts behind each study [10][11]. The field experiments, the author notes, ran in real workplaces instead of asking developers whether they felt more productive [13].

METR has since hedged its own number. In February 2026 the organisation reported that a newer experiment was affected by selection effects, concluded that its newer data was not reliable enough to estimate the current productivity effect precisely, and said AI may be speeding developers up more in early 2026 than its earlier study estimated [8][9].

So the defensible sentence carries a condition. A team of less experienced developers on repetitive work has a measured gain behind it [2][3]. A team of senior engineers inside a codebase they already know has a measured loss behind it [4][5]. The field experiments enrolled roughly 304 times as many developers as METR did, and sample size does not tell you which of the two populations is yours [17].

What to watch

  • A METR estimate for 2026 tooling that survives the selection effects it reported in February 2026.
  • Publication of how the field experiments at Microsoft, Accenture and the Fortune 100 firm defined a completed task, and the breakdown of the 26.08% by experience level.
  • A workplace experiment that measures time per task in mature codebases instead of counting tasks completed.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories