Build1 distinct publisher3 min readUpdated
The forecast implies efficiency gains get eaten by bigger models and longer agent chains. Per-agent unit economics is the line item that decides whether a pilot survives finance.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
The forecast implies efficiency gains get eaten by bigger models and longer agent chains. Per-agent unit economics is the line item that decides whether a pilot survives finance.
Research firm Gartner predicts that AI inference costs per agentic workflow will increase more than fivefold through 2028 [1]. That inverts the assumption most procurement teams are working from: token prices fall, and the bill for one completed unit of work still goes up [2].
The mechanism Gartner describes is not waste. Better efficiency lets labs develop and deploy more powerful and more expensive models, and users then find more sophisticated applications for them, agentic workflows among them, so token consumption keeps climbing [3]. Gartner calls this the inference paradox: unit economics improve while overall AI costs rise, with no clear or predictable path to matching value [4]. The savings are real and they get spent immediately, on a bigger model and a longer chain.
The arithmetic is worth doing before the next budget cycle. If the fivefold rise plays out over roughly three years, that is about 71 percent compound annual growth in the cost of running the same agent [5]. Held against a flat inference budget, a fivefold per-workflow cost means running 80 percent fewer workflows [6]. Held against a constant return, it means each workflow has to deliver at least five times the value it delivers now [7]. Gartner's framing already concedes that returns from agents are neither predictable nor guaranteed at present, and that they need to be higher than ever to justify the cost [8].
The pull in the other direction is structural. Companies are encouraged to buy agents rather than chatbots precisely because agents are more likely to deliver returns and step changes, even though agents consume many more tokens [9]. Scott Bickley, Advisory Fellow at Info-Tech Research Group, told The Deep View that the current environment has created "a top-down fervor, in fact a mandate, for virtually all enterprises to aggressively adopt AI en masse" [10]. He added that this "blind foray into the AI abyss often lacks the in-depth understanding of the total cost of ownership" and makes it hard to argue for anything but an early adopter position [11]. Bickley's view is that taking it slow would be most beneficial, and that outside pressure will not allow it [12].
The supply side is not neutral on any of this. Inference demand has become the primary focus and biggest revenue generator for many leading labs, including OpenAI and Anthropic, which are scrambling for compute to serve it [13]. Nvidia said it would provide up to 105 billion dollars in financing for OpenAI's data center in Ohio, and wrote that AI factories are the "defining infrastructure" of the era and that "compute is revenue" [14]. Compute constraints have emerged since AI was widely adopted and been made worse by growth on the research side, and both compute and inference costs rest on finite physical resources [15].
What to watch is measurement, not model selection. Ask vendors for tokens per completed workflow, not per call, and ask whether that number is trending up as chains lengthen; the cost curve lives in chain length and retry behaviour, not the price card. The same discipline applies to the tool sprawl question the newsletter raises about AI coding startups now carrying multi-billion-dollar valuations: how many of these tools a company keeps paying for once the experimentation phase ends [16]. Also watch for a stated base year on the fivefold figure. Without one, the forecast can be read as anything from a steep three-year climb to a gentler one, and the difference decides whether 2026 budgets are already short.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Gartner predicts AI inference costs per agentic workflow will increase more than fivefold through 2028.
Everyone knows AI costs are rising; what fewer expected is that efficiency gains themselves may be driving the bill higher.
According to Gartner, improved efficiency lets research labs develop and deploy more powerful and more expensive models, and users find increasingly sophisticated applications for them such as agentic workflows, so token consumption keeps climbing, driving up overall inference costs.
The report's subject is an inference paradox: even as unit economics improve, overall AI costs continue to rise without a clear or predictable path to matching value.
Returns from these investments are not predictable or guaranteed at the moment and need to be higher than ever to justify the cost.
The trap is that while AI agents consume many more tokens, companies are encouraged to invest in advanced assistants rather than AI chatbots because agents are more likely to deliver returns and make step changes in organizations.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Single-source relay of an unlinked analyst forecast
The load-bearing claim is a Gartner prediction reported by one newsletter with no link to the report, no baseline cost per workflow, no start year, no methodology, and no scenario range. Supporting material is qualitative expert opinion and secondhand relays of other outlets' reporting. Nothing in the cluster can be independently checked or falsified.
Strong capital signals, no workload evidence
Adoption evidence exists but sits one layer away from the claim: infrastructure financing at up to $105 billion, labs describing inference demand as their biggest revenue source, and large coding-tool rounds including one $500 million annualized run rate. There is no measurement of the thing being forecast - agentic workflow counts, token volumes per workflow, or enterprise inference spend - so per-workflow cost trajectories remain unobserved.
Headline figure outruns the shown evidence
A precise-sounding multiple ('5x by 2028') is presented as settled while the cluster contains no baseline, methodology, or second source, and no counter-evidence on falling per-token prices is weighed. The overstatement is moderate rather than severe because the source itself hedges on returns, quotes a caution about total cost of ownership, and frames the paradox rather than promising a cost apocalypse.
Every quoted party sells into the concern
The forecast comes from a research firm that sells subscriptions on cost anxiety, the corroborating quote comes from an advisory firm that sells total-cost-of-ownership guidance, and the supporting infrastructure claim quotes Nvidia asserting that 'compute is revenue' while committing financing that expands its own demand. The issue also carries a sponsored dataset placement alongside the analysis. None of these interests are disclosed in the text.
Directionally plausible, numerically unverified
The mechanism - efficiency gains being consumed by larger models and longer agent chains - is coherent and consistent with the compute-scarcity signals in the same source, so the direction of travel is credible. The specific multiple, timing, and any implication for a given enterprise's budget rest on one unlinked forecast from a single publisher, which caps confidence well below the midpoint.
leadership
Slack Code makes the chat window a coding surface, and a platform call for engineering leaders1 distinct publisher
product
Wu says Cognition is not for sale. The more useful fact is who bought Cursor last week.1 distinct publisher
leadership
Serval's Catalyst mines the ticket queue for automation work, not the project backlog1 distinct publisher
build
Claude Code now outruns Copilot roughly two to one in JetBrains' survey of 15,000 developers1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 18, 2026