Skip to content

Build1 publisher3 min readPublished

CTOs say AI demand is wiping out cloud CPU spot discounts of up to 90%

CTOs and infrastructure heads told the Pragmatic Engineer that cloud CPU spot capacity, once up to 90% off, is nearly impossible to get as AI absorbs supply. At standard prices, work budgeted at the deepest spot rate costs ten times as much.

The Engineer · Build desk

Illustration accompanying CTOs say AI demand is wiping out cloud CPU spot discounts of up to 90%

What happened

  • CTOs and infrastructure heads at a dinner said cloud CPU spot capacity is now nearly impossible to get without long-running relationships with the providers.
  • Reserving specific CPUs now has to happen months ahead, and providers are turning down some reservations for lack of CPUs or of the right CPU type.
  • Katelyn Lesse of Claude Platform wrote that CPUs compete with GPUs for TSMC production lines while DRAM competes with HBM for memory wafers.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • cost Any batch budget built on the full spot discount grows tenfold at standard prices, and the team that planned around spot absorbs the overrun.
  • constraint Adding CPU capacity on short notice stops being dependable when reservations need months of lead time and a provider can simply refuse them.
  • exposure Buyers without long-running provider relationships are the ones shut out of spot first, going by what the dinner attendees described.

Spot capacity is how a cloud provider sells machines nobody else is renting. The Pragmatic Engineer describes providers cutting CPU prices by up to 90% below standard for machines that were "lying dormant and unused" [1]. The discount exists only while the idle pool does. According to the newsletter, spot pricing has vanished because there is no longer any shortage of demand for CPUs [4].

The new demand is the CPU half of AI work. turbopuffer runs on CPUs in AWS, GCP and Azure [5]. "Getting CPUs is not easy anymore," its CEO, Simon Eskildsen, told the newsletter [6]. He traced the load to reinforcement learning: "During RL, they need to teach the models how to do things, like searching, and then they need the model to run software, which then takes CPUs to run." [15] Agents add more. "So as the demand curve is shifting to general purpose agents, CPU demand is also going up," he said [14]. The newsletter also points to an Uber chart of agent requests growing over the past six months [12].

Katelyn Lesse, Head of Platform Engineering for Claude Platform, wrote the supply-side account [9]. It is careful work, because each constraint is named and each can be checked. At TSMC, GPUs compete with CPUs for production lines. At SK Hynix, Samsung and Micron, HBM competes with regular DRAM for wafers [9]. AMD has no fabs, so its CPUs come out of TSMC's constrained allocation. Intel has fabs, but it has had yield problems and is moving some PC chip capacity to server chips. CPUs also need DRAM, and DRAM costs more now that memory production has shifted toward HBM [11]. "What we've ended up with is CPUs getting squeezed from both sides," Lesse wrote [10].

At the deepest discount, a spot CPU hour cost 10% of the standard price. Moving the same hours to standard pricing multiplies that line by ten [1]. Workloads that got a shallower discount see a smaller multiple. The usual hedge is a reservation, with the months of lead time the dinner attendees described [3]. The inference provider in the newsletter has the rarer problem of cash it cannot spend. A VP of Engineering there said the company is at its cloud providers' limit on GPU and CPU capacity and has been told no more is available, despite offering to take the longest leases [8].

The evidence is a dinner of CTOs and infrastructure heads, one named CEO, one unnamed VP and Lesse's written analysis [2][5][8][9]. The newsletter does not include spot price histories, regions or instance families. For the report to apply to a particular fleet, that fleet has to compete for the same CPU types in the same regions as the labs and inference providers. In my view, a CPU budget for the next few quarters should use the standard rate as its baseline and count any spot capacity as a saving when it appears. Lesse cited analysts who expect CPU supply to regain headroom before memory does, "but their expectation is that it's still going to be multiple quarters away" [13]. Eskildsen expects it to get worse first: "I would assume that it gets a lot worse before it gets better on the CPU side." [7]

What to watch

  • Whether AWS, GCP or Azure publish spot price or interruption data that shows which CPU instance families and regions have actually run dry.
  • Whether Intel's move of PC capacity to server chips shortens the multiple-quarter wait for CPU headroom that Lesse's analysts expect.
  • Whether reservation refusals reach ordinary enterprise accounts, beyond the large inference provider described in the newsletter.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories