Skip to content

Build1 publisher3 min readPublished

Gartner: agent inference cost rises 5x by 2028, so budget per workflow, not per model

The forecast implies efficiency gains get eaten by bigger models and longer agent chains. Per-agent unit economics is the line item that decides whether a pilot survives finance.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • Gartner predicts AI inference costs per agentic workflow will increase more than fivefold through 2028.
  • Everyone knows AI costs are rising; what fewer expected is that efficiency gains themselves may be driving the bill higher.
  • According to Gartner, improved efficiency lets research labs develop and deploy more powerful and more expensive models, and users find increasingly sophisticated applications for them such as agentic workflows, so token consumption keeps climbing, driving up overall inference costs.
  • The report's subject is an inference paradox: even as unit economics improve, overall AI costs continue to rise without a clear or predictable path to matching value.
  • A more-than-fivefold rise spread over roughly three years is about 71 percent compound annual growth.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

Research firm Gartner predicts that AI inference costs per agentic workflow will increase more than fivefold through 2028 [1]. That inverts the assumption most procurement teams are working from: token prices fall, and the bill for one completed unit of work still goes up [2].

The mechanism Gartner describes is not waste. Better efficiency lets labs develop and deploy more powerful and more expensive models, and users then find more sophisticated applications for them, agentic workflows among them, so token consumption keeps climbing [3]. Gartner calls this the inference paradox: unit economics improve while overall AI costs rise, with no clear or predictable path to matching value [4]. The savings are real and they get spent immediately, on a bigger model and a longer chain.

The arithmetic is worth doing before the next budget cycle. If the fivefold rise plays out over roughly three years, that is about 71 percent compound annual growth in the cost of running the same agent [5]. Held against a flat inference budget, a fivefold per-workflow cost means running 80 percent fewer workflows [6]. Held against a constant return, it means each workflow has to deliver at least five times the value it delivers now [7]. Gartner's framing already concedes that returns from agents are neither predictable nor guaranteed at present, and that they need to be higher than ever to justify the cost [8].

The pull in the other direction is structural. Companies are encouraged to buy agents rather than chatbots precisely because agents are more likely to deliver returns and step changes, even though agents consume many more tokens [9]. Scott Bickley, Advisory Fellow at Info-Tech Research Group, told The Deep View that the current environment has created "a top-down fervor, in fact a mandate, for virtually all enterprises to aggressively adopt AI en masse" [10]. He added that this "blind foray into the AI abyss often lacks the in-depth understanding of the total cost of ownership" and makes it hard to argue for anything but an early adopter position [11]. Bickley's view is that taking it slow would be most beneficial, and that outside pressure will not allow it [12].

The supply side is not neutral on any of this. Inference demand has become the primary focus and biggest revenue generator for many leading labs, including OpenAI and Anthropic, which are scrambling for compute to serve it [13]. Nvidia said it would provide up to 105 billion dollars in financing for OpenAI's data center in Ohio, and wrote that AI factories are the "defining infrastructure" of the era and that "compute is revenue" [14]. Compute constraints have emerged since AI was widely adopted and been made worse by growth on the research side, and both compute and inference costs rest on finite physical resources [15].

What to watch is measurement, not model selection. Ask vendors for tokens per completed workflow, not per call, and ask whether that number is trending up as chains lengthen; the cost curve lives in chain length and retry behaviour, not the price card. The same discipline applies to the tool sprawl question the newsletter raises about AI coding startups now carrying multi-billion-dollar valuations: how many of these tools a company keeps paying for once the experimentation phase ends [16]. Also watch for a stated base year on the fivefold figure. Without one, the forecast can be read as anything from a steep three-year climb to a gentler one, and the difference decides whether 2026 budgets are already short.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories