Skip to content

Product1 publisher3 min readPublished

Per-token AI pricing punishes the agents that succeed, SiliconANGLE column argues, citing a sponsored Futurum report

Futurum's report, sponsored by neocloud QumulusAI, says agentic tasks can burn 10 to 100 times the tokens of a single inference call. Because the meter bills every step, the agents that win real users are the ones that run through their budgets first.

The Product Desk · Product desk

Illustration accompanying Per-token AI pricing punishes the agents that succeed, SiliconANGLE column argues, citing a sponsored Futurum report

What happened

  • One organization budgeted $1 million for a year of AI work and spent it in three months because the project was so successful, according to the SiliconANGLE column's author.
  • Futurum forecasts that agent and reasoning inference will grow 219% this year, with total inference spending rising from $120 billion in 2025 to $885 billion by 2030.
  • In Futurum's survey of 824 AI decision-makers, reserved and owned infrastructure made up 66% of AI compute consumption, against 19% for on-demand cloud.
  • Amberd.ai chief executive Mazda Marvasti said some customers abandoned internally built automation tools because they could not forecast or justify the costs.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • constraint A pilot's token spend is a poor basis for a production budget, because an agent adds tokens with each step of a task while adoption adds more tasks.
  • exposure A working internal tool can be cancelled for being unforecastable, so the owner of a successful agent now carries a finance risk the pilot never did.
  • decision Each agent workload now needs a call on whether it is steady enough to move off the meter to reserved or owned capacity, the same sorting cloud teams went through.
  • contradiction The survey's heavy owned-and-reserved share looks like a verdict for committed capacity, but the column's author attributes much of it to GPU buying during shortages.

A developer calls an API after lunch and has a working prototype by the end of the afternoon, with no capacity plan and no procurement cycle [4]. That afternoon is what per-token pricing sells. The same meter then bills a production system serving 20,000 employees on identical terms [4].

What teams tell themselves users do is what a chatbot does: ask a question, get an answer [5]. What an agent does with the same request is plan, call tools, check its work, retry, hand off and summarize, generating tokens at every step [5]. Futurum's multiplier applies per task [2]. Adoption multiplies it again, once for every colleague who hears the tool works. The column's author says CIOs and CFOs raise the problem with increasing frequency, and that the most successful AI projects end up costing the most [17]. The organization that burned its $1 million in a quarter was running at $4 million a year, four times its plan [1].

Mazda Marvasti, co-founder and chief executive of Amberd.ai, described the point at which finance notices. "When they start deploying it throughout the organization, the cost starts skyrocketing because it's a useful tool that somebody built, but it's now priced on a variable basis," he said [7]. Then comes the scrutiny: "It starts getting the attention of the CFO and the CIO in terms of how much I'm exactly spending to run this tool, and whether it's worth it" [8].

A rising token count reaches finance as a bill. Usage growth is a cost figure until someone prices a completed task. Futurum's spending forecast implies inference outlays growing about 7.4-fold between 2025 and 2030 [2]. On a price that scales linearly with consumption, the column's author argues, that growth produces a budget line that outpaces the value it creates [18].

Marvasti's company is the report's showcase, and it runs on the sponsor's hardware [1][13]. Amberd.ai splits an eight-GPU Nvidia H200 server on QumulusAI bare metal into four two-GPU environments and tiers customers by latency tolerance [13]. "With one 8x H200 server, two customers pay for the entire server, and I can probably have about 30 to 35 customers running on that one server," Marvasti said [14]. The column's author cautions that the lesson is not that bare metal is cheap, and frames the economics around utilization [15]. Futurum's survey found 59% of respondents primarily run AI workloads outside hyperscaler public clouds [11].

Each agent workload sorts on two questions: whether its volume is steady enough to forecast a month ahead, and whether the team can state what one completed task is worth. Unsteady volume with unknown value is still a pilot, and per-token pricing under a hard cap suits it. Known value on unsteady volume stays on the meter, reported to finance as cost per completed task so the CFO sees a unit price instead of a growing total.

Steady volume with unknown value is the dangerous cell. It is where popular tools get cancelled on cost [9], so measurement has to come before any capacity commitment. Where both are known, the workload is a candidate for reserved or owned capacity, the choice the column says enterprises now face [12]. I would keep an agent on the meter until the team can answer both questions. The tradeoff sits in the last cell: a reserved server is a fixed cost, and it pays off only while it stays busy [15].

What to watch

  • An independent measurement of agent token use per task, from a party that does not sell compute, to test Futurum's 10 to 100 times range.
  • Whether owned and reserved capacity keeps its share of AI compute as GPU supply loosens and shortage-era purchases age out.
  • Per-task or committed-use pricing from model vendors for agent workloads, a change that would alter the pilot-to-production cost curve.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories