Leadership1 distinct publisher3 min readPublished
Enterprise AI productivity claims sit anywhere between a third off delivery time and doubled output. A randomised trial suggests much of that range comes from asking developers how fast the work felt.
The Board Room · Leadership desk

product
Alice raised $140m to red-team the frontier, and a security vendor bought in quietly1 distinct publisher
build
The best grade for controlling in-house AI agents is a C+, and buyers can now cite it2 distinct publishers
security
Frontier labs put their best vulnerability-hunting models behind vetted-defender lists1 distinct publisher
leadership
As AI budgets roughly double, repeated computation quietly adds to the bill1 distinct publisher
Compiled by The Board RoomSomething wrong?How this is made
Forty-four percentage points separate what METR's participants predicted from what measurement recorded: they expected to be roughly 25% faster and were roughly 19% slower [4][5][13]. Either end of that is interesting, but the width is the part a budget owner should hold onto. Sixteen experienced developers, working in codebases they already knew, misjudged their own throughput by that margin and were still misjudging it after the work was done [3][4]. Any productivity figure whose underlying instrument is how the work felt inherits an error bar wide enough to swallow the result.
That is the case for reading the spread in enterprise claims as an artefact of instruments rather than a genuine range of outcomes. Convert the two figures cited by Laks Alagappan, a Genpact VP who leads delivery for Tier 1 banks in the UK and Europe, into a common unit: a third off elapsed development time is about 1.5 times the prior rate, while doubled output is 2.0 times [1][2][14]. Two firms can both be reporting in good faith and land that far apart, because one is timing a coding task and the other is counting work that reached production.
The mechanism behind the disappointment is that relieving one stage does not relieve the system. Google's 2025 DORA report associated AI adoption with higher delivery throughput and lower delivery stability [7]. Alagappan reports that banking technology leaders see individual productivity rise and dashboards improve while release cadence barely moves, because change approvals, regression testing, compliance reviews, environment provisioning and operational readiness still set the pace [9]. Once code generation stops being the binding constraint, the queue forms at whatever comes next [12].
The strongest objection to this reading is that sixteen open-source developers on 2025 tooling is not an evidence base for a capex decision, and that METR's own 2026 follow-up suggested newer tools were likely delivering larger gains, with selection effects leaving the size unreliable [6]. That objection is right on its own terms: the 19% figure should not be used as an enterprise baseline. The durable finding is narrower. Self-assessment did not track measured output, which disqualifies the developer survey as a primary instrument whatever the true effect from current tools turns out to be, and on that size the record says we do not know yet.
The decision available this quarter is which instrument to fund, not which tool to buy. Coding speed, lines of code, suggestion acceptance rates and individual output are cheap to collect and, in Alagappan's reading, are not measures of engineering productivity at all [8]. DORA's observation that AI's benefits scale with the maturity of the surrounding system points spend toward automated testing and fast feedback loops [10], and the pre-code work Alagappan flags, including de-conflicting requirements, assessing regulatory impact, mapping application dependencies and investigating production incidents, is where he expects the larger return while being harder to put on a dashboard [11]. The ordering consequence is concrete: a team that starts reporting delivery outcomes now will be able to defend or kill this spend a year from now, and a team reporting acceptance rates will be asked why cadence did not change and will have collected nothing that answers it.
Ranked by verification strength, evidence, and original report placement.
Laks Alagappan, author of the Forbes Tech Council piece, is VP and Client Engagement Manager at Genpact and leads enterprise technology delivery for Tier 1 banks across the UK and Europe.
In 2025 the AI safety research nonprofit METR conducted a randomized controlled trial with 16 experienced open-source developers working on mature codebases they knew well.
Participants in the METR trial expected AI to make them about 25% faster and still believed it had afterwards.
Objective measurement in the METR trial showed participants were about 19% slower.
Google's 2025 DORA State of AI-Assisted Software Development report found that AI adoption was associated with higher software delivery throughput but lower delivery stability.
Alagappan writes that many organizations still evaluate AI using local metrics such as coding speed, lines of code generated, AI suggestion acceptance rates and individual developer output, and that these are not measures of software engineering productivity.
Distinct publishers with included, body-backed reporting in this cluster.
forbes.com
1 article · September 3, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Two studies, quoted from memory
Every load-carrying number reaches the reader through one Forbes Tech Council post: the 25% expectation, the 19% slowdown, the throughput-up-stability-down pairing. METR's trial is simultaneously the strongest thing in the story and its thinnest sourcing — sixteen developers, no design detail, no link, and a 2026 follow-up that points the other way compressed into a single hedged sentence. The DORA findings arrive as paraphrase with no figures attached.
Nothing counted
The piece tells us dashboards look encouraging while release cadence barely moves, and attributes that to unnamed banking technology leaders. That is the shape of an anecdote, not a measurement: no seats, no deployments, no release frequency, no before-and-after. We decline to score adoption from it.
Anti-hype piece, overextended
The column argues against inflated claims, which makes its own reach easy to miss. Sixteen open-source developers on codebases they already knew become a verdict on measurement inside Tier 1 banks. The perception gap is genuine and worth the attention; the generalisation runs ahead of what one small trial can carry, and the evidence that newer tools do better is waved past in a sentence because it complicates the frame.
Services byline, no editor in between
The punchline — that the value lies in requirements, testing, governance and operations rather than code generation — is a description of what Genpact sells to the banks the author covers. Forbes Councils is a paid membership platform where members publish themselves, so nothing stood between the thesis and the page. None of that makes the METR figures wrong; it does mean the framing arrived pre-aligned with the writer's book of business, and the piece never says so.
Direction firm, particulars loose
We hold the core reasonably firmly — a measured slowdown alongside a felt speedup is a well-travelled result, and the systems argument built on it is coherent. Everything stacked above it is looser: an unquantified 2026 follow-up, unnamed bank leaders, DORA without numbers, one publisher, one interested author, and no second account to check against.