Leadership1 distinct publisher3 min readUpdated
AI is billed by usage, and the introductory pricing has ended. The programs in trouble are the ones that put exhaustive, repeatable checking work onto a metered model endpoint.
The Board Room · Leadership desk
Compiled by The Board RoomSomething wrong?How this is made
The gap inside IBM's finding is the part worth taking to a budget meeting: 53 percentage points between the projected return and the actual one, on the same project [10]. No productivity claim being made for AI tooling is that large. A business case with an error bar that wide is not a forecast, it is a hope with a spreadsheet attached. And note which word is load-bearing in IBM's framing: unaccounted [9]. The debt was always there. The accounting was not.
Why the spend fails to convert is a question about how the tools behave. Large language models are trained to be helpful and generative, which means they are not optimized for exhaustive validation [13]. Asked to review a codebase, a model does not read every line; it searches for areas of interest and evaluates the context around them [14]. It leans toward returning the first good answer, and has to be explicitly told or configured to do more thorough work, which is where the iteration cost lands [15]. So a compliance-shaped question, of the form "does every element meet the standard," gets answered by a system that samples. More money buys more samples. It does not buy the guarantee.
That sets the sorting rule. Where the requirement is checking every element against a defined standard, every time, without exception, deterministic rules-based tools beat model calls on speed, cost, consistency and completeness [16]. The model earns its keep on the judgment calls, the synthesis, and the explanation of what a result means and how to fix it [17]. The criterion for which side a task belongs on is the requirement, not the department that owns it, and most AI programs have never run that sort.
The argument arrives from an interested party, which is worth saying plainly. It comes from the CTO of Deque Systems, writing from digital accessibility [1], a field he chooses as evidence because it has codified standards and measurable outcomes, so the difference between what AI does well and what it does badly shows up in numbers [19]. Take the recommendation with that in mind. The third-party figures he cites are checkable, and the description of model behaviour under a thoroughness requirement is not controversial.
His proposed remedy is unglamorous: an analytical framework built on consistent measurement of cost, speed and quality [20]. It reads like housekeeping until you remember what changed underneath it. AI is not priced by seat, it is priced by consumption [3], and the introductory rates were set to drive adoption rather than to reflect what serving the model actually cost [4]. Under seat pricing, an unmeasured task is a rounding error. Under consumption pricing, an unmeasured task is an open invoice, and the earliest adopters wrote their ROI cases in a currency that has since been repriced [2].
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
The article is written by the CTO at Deque Systems, author of the "Agile Accessibility Handbook, A Practical Guide to Accessible Software Development at Scale."
AI is not priced by seat; it is priced by consumption.
Early on, frontier-model companies set prices that did not reflect actual cost, in order to drive adoption and further develop their models, which led to widespread adoption in which effectiveness and ROI were secondary to whether the problem could be solved with AI at all.
Uber exhausted its entire annual AI coding budget before summer, and COO Andrew Macdonald said publicly that the "link is not there yet" between token consumption and useful products shipped.
According to the FinOps Foundation, AI cost management was a concern for roughly one-third of financial operations practitioners in 2024, and by 2026 it concerns nearly all of them.
Per IDC, as reported by CIO, Global 1,000 companies will underestimate their AI infrastructure costs by 30% through 2027.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Thin: one vendor-authored contributor post relaying unlinked third-party figures
The cluster contains exactly one source, a Forbes Tech Council contributor essay written by the CTO of a company selling the recommended remedy. Its factual spine (Uber's exhausted budget and COO quote, FinOps Foundation concern trajectory, IDC via CIO, IBM IBV, WebAIM, CodeRabbit, Deque's own survey) is attributed but never linked or dated, one central anecdote is an explicitly unnamed rumor, and the strongest technical assertion has no published measurement behind it.
Real cost-governance pressure indicated, but all signals are second-hand
There are concrete adoption-adjacent datapoints — a named enterprise (Uber) exhausting an annual AI coding budget with an on-record executive quote, a described shift in FinOps practitioner concern toward near-universal, an IDC cost-underestimation forecast, plus WebAIM, CodeRabbit and Deque survey measurements. That is more than pure assertion, but nothing is verifiable inside the cluster, one figure is anonymous rumor, and one survey is run by the author's employer, so measured adoption stays low.
Overstated: strong causal framing on anecdote, rumor, and vendor-aligned inference
The headline proposition that the cheap-token era has ended and consumption bills are outrunning business cases is asserted without a single pricing datapoint, while the remedy — deterministic rules engines over metered model calls — is exactly what the author's employer sells. The governance concern is genuine and echoed by the cited FinOps and IDC material, so the gap is moderate rather than extreme, but the certainty of the framing outruns what this cluster demonstrates.
High: vendor CTO in a paid contributor channel arguing for his own product category
The author is CTO of Deque Systems, a deterministic digital-accessibility testing vendor, publishing through Forbes' council contributor program, citing his own company's 2026 survey as market data, concluding that deterministic rules-based tooling beats AI on exhaustive verification, and closing with a teaser for a follow-up installment. Affiliation is disclosed, which limits the concern somewhat, but the alignment between conclusion and commercial interest is direct.
Low: directionally credible governance thesis, weakly sourced specifics
Confidence is limited by single-source, single-author coverage with high incentive alignment and no linkable primary evidence. The general mechanism — consumption pricing plus absent usage limits produces budget shocks, and exhaustive repeatable checking is a poor fit for metered token endpoints — is coherent and partly corroborated by attributed third-party research, so the story is not dismissible; the specific numbers and the categorical pricing claim should not be relied on without independent verification.
build
AI-written code fails the same four ways, and every gate you own reports green1 distinct publisher
leadership
Robotaxis Are Taking Mid-Teens Share in Three Metros. Headcount Will Not Show It.1 distinct publisher
product
Uber's first European robotaxi still has a driver in it, and that is the whole story1 distinct publisher
product
APIs built for human judgment now answer to agents that have none1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 21, 2026