Skip to content

Build1 publisher2 min readPublished

Anthropic survey: engineers' merged PRs rose 67 percent, but org's delivery dashboard reportedly unmoved

Anthropic's own engineers report large gains from Claude Code while the delivery dashboard sits still. The dev.to post that collected the numbers argues that both of the dashboards most teams watch measure only cost.

The Engineer · Build desk

Illustration accompanying Anthropic survey: engineers' merged PRs rose 67 percent, but org's delivery dashboard reportedly unmoved

What happened

  • Anthropic surveyed 132 of its own engineers about Claude Code and found merged pull requests per day up 67 percent.
  • Daily use of the tool among those engineers climbed from 28 percent to 59 percent, with self-reported productivity gains of 20 to 50 percent.
  • When the organization's delivery dashboard was checked, the delivery metrics were flat.
  • McKinsey found 30 percent of leaders could say where the time AI freed up actually went.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint A team whose only instruments are an adoption chart and a token meter has priced its spend and measured nothing on the other side; neither instrument follows freed hours to an outcome.
  • decision Anyone writing an AI business case now has to choose whether to split claimed value into hard and soft dollars, knowing the hard total is the only half a CFO can check against a budget that stayed the same.
  • exposure Engineering cases are the most exposed to challenge, because a three-weeks-to-one-week claim dollarizes queue and review waits that no engineer worked.
  • contradiction Self-reported gains of 20 to 50 percent and a flat delivery dashboard cannot both be pricing the same thing, so a buyer weighing the tool is really choosing which boundary to measure at.

Merged pull requests per day is counted at the developer's edge of the pipeline. Delivery metrics are counted at the customer's. Between them sits what the dev.to post calls cycle time, the calendar span from request to delivery, and the post breaks that span into a ticket sitting in a queue, a wait on review, and a dependency that had not shipped [14]. None of those is labor time. The post's advice for that case is to report the improvement as a rate, this class of work now moves 40 percent faster, and to call it money only when shipping sooner brings revenue in sooner [15].

Adoption tells you whether anyone is using the thing, and the post grants that it is a useful leading indicator [8]. Token spend tells you what the tool costs to run, and it belongs on the cost side of the calculation [9]. Both come out of a dashboard with one query. The post says that is exactly why they are the two numbers that get reported [16].

The two survey figures also move at different speeds and in different units. The share of engineers using Claude Code daily rose 111 percent on its base, against 67 percent for merged PRs per day [1]. One is a share of people and the other is a rate of output, so their ratio tells you nothing about efficiency. It does show that the output metric grew more slowly than the adoption metric, and that neither of them is a delivery metric. The post does not say which delivery metrics were checked, or over what window [5].

The business case the post describes multiplies saved hours by hourly cost and reports the total as money saved [10]. The objection is a budget objection. Nobody was let go and no contractor was dropped, so the company pays the same salaries to people with slightly lighter weeks [11]. Hours become money in three cases: the time is redeployed onto work that generates value, it avoids a hire you were about to make, or the same headcount produces more of something you sell [12]. The remedy the post proposes is to label every claimed dollar hard or soft and report the two totals separately [13]. I would take that split over any adoption chart.

The wider figures come from surveys of leaders. Gartner's 2025 number puts the leaders who said their AI tools had returned significant value at 22 percent, and the other 78 percent did not say so [2]. The post says that 22 percent lands where McKinsey, Deloitte and ServiceNow each arrived measuring it their own way [7].

What to watch

  • Whether Anthropic publishes the delivery metrics and the window behind the flat line, which would show where in the pipeline the 67 percent stopped.
  • Whether Gartner's next read moves off 22 percent of leaders reporting significant value.
  • Whether any coding-assistant vendor ships hard and soft dollar reporting to replace the adoption and token charts.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories