Leadership1 distinct publisher3 min readPublished
Anna Meadows, writing for the Forbes Tech Council, cites a study of roughly 800 developers producing 41% more bugs with no throughput gain. Review was never sized for the new output.
The Board Room · Leadership desk

Compiled by The Board RoomSomething wrong?How this is made
The mechanism worth dwelling on is not volume, it is plausibility. Meadows argues that machine-written code rarely looks lazy; it looks idiomatic and finished right up to the point where it is quietly wrong about your business logic, and no linter catches that [4]. Human review heuristics were tuned on human failure signals. Those signals are absent, so the reviewer's instinct returns a false negative and the approval lands [9].
Her two illustrative figures, a 900-line pull request [9] and the 90 seconds of scrutiny a uniform pipeline gives every change [10], sit in different paragraphs. Put them together and the implied reading rate is ten lines a second [14]. That is rhetoric, not measurement, but it names the gap the velocity chart hides.
The quantitative load in the piece rests on one study: roughly 800 developers, 41% more bugs from AI-assisted engineers, no meaningful throughput gain [2]. The column does not name it, its authors, or where it was published [c2b], which is worth knowing before anyone puts it in a board deck. If it holds, the accounting is worse than the headline. Bug count up 41% with output flat means defects per unit of shipped work also rise 41% [13]. And the column opens with the opposite feeling: teams adopting these tools, watching pull requests get bigger, and celebrating [15]. Bigger diffs are not more delivered value, and leaders should be clear about which of those two their dashboards actually count.
The three prescriptions all move work earlier, and each has a price someone has to approve. Making tests the specification means the engineer writes or signs off the intent before the generator runs [7], which spends exactly the minutes that felt like the speedup. Deterministic gates, contract tests, property-based tests, static analysis, policy checks that block known-bad patterns, are engineering projects with owners and maintenance, not settings to enable [6]. Risk-routed review requires somebody to declare, in advance and in writing, that billing logic and a dependency bump do not travel the same path [8]. That declaration is a governance artifact, and most teams do not have one.
Then the part that outlives the incident. Meadows' answer to who is accountable for agent-written code is whoever merged it, but only if the process makes that true: you need a durable record of what was generated, what was modified, what was checked, and who made the call [11]. Today that context lives in chat threads, tool sessions and people's heads, and it evaporates within weeks, well before an auditor or a postmortem asks [12]. So the practical question for anyone running this stack is a retention question. Does the merge record say which lines a model produced and which gates ran, or does it say "LGTM"?
Teams that scaled generation without scaling verification have already taken on the liability [5]. The invoice arrives as incidents, rework, and code with a human name on the commit and no human who can explain it [5].
Ranked by verification strength, evidence, and original report placement.
First prescribed shift: move trust from reading to checking, with deterministic gates carrying as much load as possible (type systems, contract tests, property-based tests, static analysis, policy checks that block known-bad patterns), and humans as the last check rather than the only one.
On accountability for agent-written code, Meadows says the instinctive answer, whoever merged it, is the right answer but only if the process makes it true, which requires knowing later what was generated, what was modified, what was checked and who made the call.
Meadows says engineering activity is surprisingly ephemeral: context lives in chat threads, tool sessions and people's heads and evaporates within weeks, before an auditor, a security team or a postmortem asks six months later.
Anna Meadows, CTO of CodeROI, writing for the Forbes Tech Council, argues that writing code used to be the expensive part of software and is not anymore; the expensive part is now verification, deciding whether the code is correct, safe, and something you want to own for the next five years.
The Forbes column cites the study of roughly 800 developers without naming it, its authors, or where it was published.
Meadows says AI tools multiply the volume of code entering review but do nothing to multiply the humans doing the reviewing, so reviewers adapt by skimming, approving, and trusting the green checkmarks.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Thin: one vendor-authored opinion column, single uncited statistic
Everything in the cluster comes from one contributed Forbes Tech Council column. Its qualitative claims are verifiable only as statements of the author's view; its single quantitative anchor -- roughly 800 developers, 41% more bugs, no throughput gain -- carries no study name, authors, venue or methodology, and no second source corroborates it. The prescriptions come with no measured outcomes.
No adoption data supplied
The column asserts anecdotally that teams adopted AI coding tools and that most organizations treat verification as an afterthought, but supplies no releases, deployments, benchmarks, usage disclosures or counts. No adoption observation can be recorded without inferring facts the source does not provide.
Overstated: data-flavored framing on one uncited number
The framing is more certain than the evidence: 'the data backs this up', 'Generation is solved. Verification is not', and a claim that regulators, insurers, acquirers and tax authorities are already asking for code provenance. Behind that sit one unattributed study, illustrative round figures (900 lines, 90 seconds) presented as if diagnostic, and no named regulation or measured outcome. The underlying observation -- review capacity did not scale with generated volume -- is plausible and specifically argued, which keeps the gap from being extreme.
Strong: vendor CTO in a paid-contributor channel prescribing her own category
The author is CTO of CodeROI, described in her own byline as 'building deterministic infrastructure for regulated software workflows', and the piece's prescriptions -- deterministic gates, captured provenance of what was generated/modified/checked, audit-ready records for regulators and insurers -- map directly onto that category. It runs in the Forbes Tech Council contributed channel, where placement is membership-driven rather than editorially commissioned, and no independent voice or counter-evidence appears.
Moderate: what was said is clear, whether it is true is not
The source text is complete and unambiguous about what is claimed, who claims it and where they sit commercially, so attribution-level claims are solid. Confidence in the substantive picture is capped by a single publisher, a single item, an unverifiable central statistic and no adoption or outcome data.
leadership
Gartner says agents aren't ready; 60% of companies plan to deploy them anyway1 distinct publisher
leadership
Enterprise security reviews went from 20 questions to hundreds of rows, and vendors pay first1 distinct publisher
leadership
Your fastest-moving insider now holds an API token, not a grudge1 distinct publisher
leadership
Physicians' 86% privacy condition makes clinical AI a processing question, not a model question1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
forbes.com
1 article · August 26, 2026