Skip to content

Product1 publisher2 min readPublished

Qualcomm puts tokens per watt at the centre of its AI infrastructure pitch

Qualcomm's Tony Pialis told the AI Infra Summit in Santa Clara that tokens per watt is the new key metric in the AI war. The vendors in the room disagreed about which memory architecture gets them there.

The Product Desk · Product desk

Illustration accompanying Qualcomm puts tokens per watt at the centre of its AI infrastructure pitch

What happened

  • While tech titans at Salesforce's Dreamforce in San Francisco debated slowing AI deployment, technologists met an hour's drive south at the AI Infra Summit in Santa Clara to describe building AI infrastructure as fast as possible.
  • Qualcomm's datacenter and AI chief Tony Pialis told the conference on Wednesday that tokens per watt has become the new key metric in what he called the AI war.
  • Amazon senior vice president Peter DeSantis said the majority of compute will serve inference, and AWS launched its Graviton5 Arm CPU in June for real-time AI reasoning and multistep task orchestration.
  • SiliconANGLE's analysts put AI server memory spend at $35 billion in 2025 and between $175 billion and $190 billion by 2027, with a typical AI server using about eight times the memory of a traditional one.
  • Qualcomm has moved off high-bandwidth memory to a proprietary high-bandwidth compute architecture, while d-Matrix has stacked 3D DRAM into its next chip architecture, Raptor.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • cost Enterprises are the ones carrying token bills that SiliconANGLE calls out of control, and the improvement the suppliers are promising is denominated in watts. Watts never show up on a customer invoice.
  • contradiction Qualcomm rates its memory architecture per watt against HBM while d-Matrix rates its stacked DRAM against SRAM, so a buyer comparing inference platforms is reading two claims measured against two different incumbents.
  • decision Anyone signing a multi-year inference commitment is taking a position on DRAM prices, because memory is the input projected to grow five-fold inside the term of that contract.
  • constraint Inference-led demand pulls CPUs back into AI capacity planning. Teams that standardised on GPU-only serving face a re-platforming question they did not budget for.

Tokens per watt is a supplier's unit. Watts are what the vendor pays for; tokens are what the product team buys. Pialis said "We clearly cannot stay on this trend. We on the infrastructure side need to do better." [4] SiliconANGLE did not report a price per token, or a date when one comes down. [19]

The memory projection is the number a planner can hold on to. Subtract the 2025 figure from the 2027 range and it adds $140 billion to $155 billion of memory spending in two years, 5.0 to 5.4 times the 2025 level [12]. Token costs are already spiraling out of control for many enterprises, according to SiliconANGLE [5].

Qualcomm and d-Matrix are both working on the memory wall, which Pialis describes as compute having outrun both memory and bandwidth [18]. Qualcomm estimates its high-bandwidth compute architecture at six times the bandwidth per watt versus HBM for large batch sizes, and 200 times the capacity per watt versus SRAM [14]. "The bottleneck is data movement, not arithmetic," Pialis said. "The way to solve that is to bring the compute even closer to the memory." [15] d-Matrix founder and chief executive Sid Sheth said of his company's stacked DRAM, "We are much better than SRAM" [17].

The six-times figure is stated for large batch sizes [14], the regime where many requests are served together. A product whose traffic arrives in bursts does not run there.

The lever a team can pull this quarter is cheaper silicon around the model. DeSantis said "Today, the vast majority of workloads on AWS run cost effectively and faster on Graviton" [9], and Graviton is AWS's family of 64-bit Arm-based CPUs built for energy efficiency in cloud applications [8]. Moving retrieval, queues and API handlers onto that hardware lowers the cost of the code around the model and leaves the token line where it is.

Two questions sort the AI features in next year's plan. Whether cost per active user rises the more that user uses it, and whether tokens can be capped without breaking the promise the feature makes. Flat-cost and cappable features ship at today's prices, whatever happens on the infrastructure side. A feature whose cost climbs with use and cannot be capped, such as an assistant that re-reads a whole workspace on every question, is being underwritten by a bet that tokens per watt improves on the vendors' schedule.

What to watch

  • Whether Qualcomm publishes HBC bandwidth-per-watt figures at small batch sizes, or independent benchmarks appear.
  • Whether AI server memory spend tracks the $175bn to $190bn range projected for 2027, since DRAM pricing ends up inside per-token rates.
  • Whether any inference provider quotes a dated price per million tokens for 2027 capacity.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories