Published · 3d agoScience2 min read
Token prices fall 95%, agent bills rise 5x: the cost lives in the orchestration graph
Gartner projects agentic inference costs rising more than fivefold in two years even as token prices collapse. The variable that moves is the runtime, not the model.
Written for builders.See today for builders

What happened
- Gartner research predicts that token costs will fall by 95% by 2030 while inference costs for agentic workflows will increase more than fivefold over the next two years, because AI app builders use more and often more expensive tokens as LLMs get more complex.
- Gartner analysts Will Sommer and Sabine Zimmerhansl call the gap the "inference paradox" and write that "the market is captured by a token-deflation illusion"; buyers assume provider token-economics savings will be reflected in their roadmaps, but "they will not."
- Gartner predicts that by 2030, performing inference on an LLM with one trillion parameters will cost GenAI providers over 90% less than in 2025, driven by semiconductor and infrastructure efficiency improvements, model design innovations, higher chip utilisation, increased use of inference-specialised silicon, and edge devices for specific use cases.
- Gartner reports that agents require 5x to 30x more tokens than a chatbot to handle equivalent tasks, that agent inference costs are 5x greater, that training hardware costs for medium-sized agentic reasoning models are 2.5x those of similarly sized chatbots, and that advanced reasoning agents already cost up to 150x more on a single task than basic AI chatbots.
- Gartner built a Tokenomics Model and ran 12 types of AI model, finding basic workflows cost around $0.05 per inference token, summarization and knowledge retrieval roughly $0.10, more complex workflows around $0.30, and planning and learning roughly $0.40 per token, making provider cost per token for planning and learning 8x to 10x that of basic workflows.
Compiled by The ScientistSomething wrong?How this is made
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
Gartner research predicts that token costs will fall by 95% by 2030 while inference costs for agentic workflows will increase more than fivefold over the next two years, because AI app builders use more and often more expensive tokens as LLMs get more complex.
- [2]
Gartner analysts Will Sommer and Sabine Zimmerhansl call the gap the "inference paradox" and write that "the market is captured by a token-deflation illusion"; buyers assume provider token-economics savings will be reflected in their roadmaps, but "they will not."
- [3]
Gartner predicts that by 2030, performing inference on an LLM with one trillion parameters will cost GenAI providers over 90% less than in 2025, driven by semiconductor and infrastructure efficiency improvements, model design innovations, higher chip utilisation, increased use of inference-specialised silicon, and edge devices for specific use cases.
- [4]
Gartner reports that agents require 5x to 30x more tokens than a chatbot to handle equivalent tasks, that agent inference costs are 5x greater, that training hardware costs for medium-sized agentic reasoning models are 2.5x those of similarly sized chatbots, and that advanced reasoning agents already cost up to 150x more on a single task than basic AI chatbots.
- [5]
Gartner built a Tokenomics Model and ran 12 types of AI model, finding basic workflows cost around $0.05 per inference token, summarization and knowledge retrieval roughly $0.10, more complex workflows around $0.30, and planning and learning roughly $0.40 per token, making provider cost per token for planning and learning 8x to 10x that of basic workflows.
- [6]
An agent that queries a CRM at step four and gets back 400 rows will, by step nineteen, have had the model read those rows fifteen more times at input rates, because the harness assembled that prompt on every turn and kept the rows in it.
Sources & coverage · 5 publishers
The reporting this story was synthesized from, earliest first. Every link goes to the original.


