Skip to content

Leadership1 publisher3 min readPublished

The scarce resource in software delivery moved from writing code to checking it

Daily agent use among engineers climbed to roughly 80% in a year while developer trust in the output fell to 29%. The case that the gap is costing delivery time rests on a randomized trial of 16 people.

The Board Room · Leadership desk

Illustration accompanying The scarce resource in software delivery moved from writing code to checking it

What happened

  • Temporal's 2026 State of Development Report, a survey of 554 engineers and engineering leaders in the US and UK, found roughly 80% now use AI agents daily or more often, up from 47.3% a year earlier.
  • Stack Overflow's most recent Developer Survey put developer trust in the accuracy of AI output at 29%, down from 40% a year earlier, with tool usage at 80%.
  • In that survey, 45% named AI output that is "almost right but not quite" as their single largest frustration, and 66% said they now spend more time fixing such code.
  • METR's randomized controlled trial put 16 experienced open-source developers on 246 real issues in repositories they knew, and measured them 19% slower when permitted to use AI tools.
  • Sonatafy's Software Delivery Failure Index 2026 analysed 195 interviews with software leaders, identified 48 documented delivery failures across five root causes, and put the largest single overspend at $42 million.

Compiled by The Board RoomSomething wrong?How this is made

Why it matters

  • constraint Weekly review throughput is the real ceiling on shippable work, and while it stays unwritten, planning built on generation capacity keeps committing work that nobody has budgeted anyone to verify.
  • contradiction Stack Overflow's 29% accuracy trust and Temporal's 85.5% trusting at least somewhat answer different questions, so how alarming the confidence picture looks depends on which instrument a leader accepts.
  • exposure A board reading throughput and deployment frequency alone cannot see rework accumulating, so the cost lands later as schedule slip or overspend, and the measures in front of the board tie it to nothing.
  • decision Accepting a team's self-reported AI productivity gain now means explaining away the one trial where skilled practitioners misjudged their own speed by dozens of points on code they knew.

Generated code enters the same pipeline as hand-written code. The extra work arrives downstream as rework, and it lands on a line item that never names the tool that produced it. Steve Taplin, CEO and founder of the software consulting firm Sonatafy Technology, calls this the AI validation gap, which he defined as "the absence of evaluation, measurement or trustworthy feedback loops on AI initiatives" [9][12]. He wrote that the same gap is harder to see inside a delivery system than in a product, because the output looks like ordinary code [10]. His framing of what changed: "For 30 years, the scarce resource in software delivery was the ability to write code," and what replaced it "is the ability to determine whether the code that now exists is correct, appropriate to the architecture and safe to release" [22].

The Temporal survey shows the split inside one population. Some 85.5% of respondents trust agent output at least somewhat, and 41.1% hit agent-related issues daily or more often [8]. Those groups cannot be separate: at least 26.6% of respondents both trust the output and are interrupted by it every day [19]. The median respondent runs five agents [2].

METR's trial is the only measurement in this set that put a self-report next to an observation. The developers had forecast 24% faster, and afterwards, having lived through the slowdown, still estimated a gain of about 20% [6]. That is a 43-point gap between forecast and result, and a 39-point gap between the post-hoc estimate and the result [20]. Sixteen developers in their own repositories cannot carry a claim about industry delivery speed, and METR says the sample is small [7]. The two surveys are self-report by construction, so they cannot correct it.

The failure index cited alongside that trial is Taplin's own firm's work, and 48 documented failures out of 195 interviews is 24.6% of the sample against criteria the firm set [11][21]. Taplin wrote that "AI did not invent a new way for software projects to fail. It removed the friction that used to slow the old ways down, faster than most organizations updated their instrumentation" [13].

None of the three datasets measures delivery cycle time in the same organisation before and after agents arrived. The claim in the article's headline, that 80% adoption has not made software delivery faster, is an inference assembled from an adoption survey, a trust survey and a 16-person trial [24]. The evidence supports something narrower, and that narrower thing is still useful: generation capacity moved a lot in twelve months, review capacity is still one senior engineer reading at roughly the speed of 2019, and two thirds of developers report more time spent fixing near-miss output [1][14][5].

The change Taplin proposes for the next planning cycle is cheap. Report change failure rate, rework rate and time to restore next to pull request throughput, because tracking throughput alone means "you have built a scoreboard that cannot report a loss" [15]. Then write down how many pull requests a senior engineer can meaningfully review in a week, a number he said leaders often shrug at while knowing headcount to the person [16].

What to watch

  • Whether Stack Overflow asks the same accuracy-trust question next year, which would turn the 40-to-29 move into a trend rather than a single-year reading.
  • A replication of METR's design with a larger sample, or one run in a commercial codebase. METR's developers worked in their own open-source repositories.
  • Whether any engineering organisation publishes change failure rate and rework rate next to its agent adoption numbers.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories