Bottom line: According to Gartner, falling token prices will be massively outweighed by the sharply higher token consumption of AI agents, causing total inference costs to rise more than fivefold by 2028.
Gartner forecasts that enterprises will pay more than five times today’s inference costs for agentic AI workflows by 2028 – even though the cost per token continues to decline. The reason is the sharply rising token consumption of complex AI agents compared to simple chatbots.
In a recent forecast, the analyst firm Gartner describes a phenomenon it calls the “inference paradox”: although the cost per processed token for large language models continues to fall, and new model generations compute more efficiently than their predecessors, enterprises will, according to Gartner, have to spend significantly more on AI inference in the coming years. Per agentic workflow, this increase is expected to amount to more than five times today’s costs by 2028. The reason is the shift from simple, assistive AI functions to AI agents that autonomously perform multi-step tasks. Gartner analyst Will Sommer explains the difference: a chatbot must read a request, interpret it, and respond with a probably fitting answer, while an AI agent must continuously draw conclusions, negotiate, and question itself. According to Gartner’s calculations, routing a task to an agentic reasoning model causes at least five times the inference costs for providers compared to a simple chatbot interaction, and for more complex tasks this factor can be significantly higher, according to Gartner.
For decision-makers in enterprises, it is relevant that the cost structure does not scale linearly with the capabilities of the models. Gartner identifies three trends shaping the so-called token economy: first, the cost structure of foundation models is improving rapidly. Second, more efficient models enable the use of even more powerful and expensive models for more demanding use cases. Third, more complex AI workflows consume considerably more tokens than simple chat interactions, driving up total costs. Tokens are thus becoming cheaper, but not quickly enough to keep pace with the growth in AI capabilities and the associated costs. The pace of innovation, accordingly, clearly outstrips the cost curve.
For product owners, this means, according to Sommer, that they cannot rely on a more efficient token economy to automatically justify AI costs. Each new generation of AI capabilities requires more, and often more expensive, tokens. A universal model that is both powerful and cheap is not in sight. Anyone who wants to build competitive AI products cannot avoid building and maintaining complex ecosystems made up of multiple models.
To achieve a positive return on investment from advanced AI capabilities such as reasoning agents despite rising costs, Gartner names two possible paths: either returns from deployment must grow exponentially along with costs, or enterprises rely on finely tuned inference tiering, routing, and orchestration to assign tasks to the most cost-efficient models depending on complexity. Both are feasible, but according to Gartner require considerable effort across entire workflows. Sommer summarizes: those who default to generic autonomous intelligence will face unbounded costs that are orders of magnitude higher than those of optimized product ecosystems.
Source: www.it-daily.net · Published August 17, 2026
Lumi AI News — AI-assisted curation pursuant to Art. 50 EU AI Act. Paraphrasing and classification by Lumi News Pipeline v1.8.3.