Skip to content

Leadership2 publishers3 min readPublished

Infrastructure leaders say AI demand has erased cheap CPU spot capacity

CTOs told Gergely Orosz that CPU spot pricing, once up to 90% below list, has nearly vanished as AI workloads absorb spare cloud capacity. Teams whose compute budgets assumed spot rates now have to plan for getting capacity at all, on top of paying more for it.

The Board Room · Leadership desk

Illustration accompanying Infrastructure leaders say AI demand has erased cheap CPU spot capacity

What happened

  • turbopuffer CEO Simon Eskildsen, whose product runs on CPUs, said reinforcement learning at AI labs and general-purpose agents are driving CPU demand up.
  • An unnamed VP of Engineering at a large inference provider said it has hit the limit of what its clouds will rent, despite cash and willingness to sign the longest leases.
  • Katelyn Lesse of Claude Platform wrote that CPUs compete with GPUs for TSMC production lines while HBM pulls memory makers' wafers away from regular DRAM.

Compiled by The Board RoomSomething wrong?How this is made

Why it matters

  • cost Teams whose budgets assumed spot rates absorb the difference, which reaches up to ten times the old price for workloads bought at the deepest discount.
  • decision Booking capacity months ahead makes teams commit to a demand forecast this quarter, so a wrong forecast next quarter means idle reserved machines or a refusal.
  • exposure Accounts without long provider relationships are most exposed, since attendees said spot CPUs are nearly impossible to get without those connections.

The board-deck version of this story is a change to one line item. Spot CPUs that customers could once buy at up to 90% below the standard price are going away, so compute costs rise [1]. At the deepest discount the move is tenfold: a machine billed at 10% of list costs ten times as much at list [1]. The 90% figure was the ceiling, so a fleet that bought at shallower discounts faces a smaller jump.

That version is incomplete because it assumes money can still buy capacity. At a dinner of CTOs and infrastructure heads, attendees told Gergely Orosz that specific CPUs now have to be reserved months in advance, and that cloud providers turn down some reservations because they lack enough CPUs or the right type [2][4]. A VP of Engineering at a large inference provider, not named in the newsletter, told Orosz the company has cash, will accept the longest leases, and is still told no more GPU or CPU capacity is available [10].

The demand account comes from buyers. Simon Eskildsen is CEO of turbopuffer, whose product runs on CPUs across AWS, GCP and Azure [5]. He tied the squeeze to reinforcement learning, where models run software during training, and to agents doing general-purpose work [9]. "Getting CPUs is not easy anymore," he said [6]. "So the labs are sucking up a lot of CPUs" [7]. He does not expect quick relief. "Even the big companies are fighting each other for the right to get the CPU allocations. I would assume that it gets a lot worse before it gets better on the CPU side," he said [8].

The supply account comes from Katelyn Lesse, Head of Platform Engineering for Claude Platform. She wrote that GPUs compete with CPUs for production lines at TSMC, and that HBM competes with regular DRAM for wafers at SK Hynix, Samsung and Micron [11]. "What we've ended up with is CPUs getting squeezed from both sides," she wrote [12]. AMD's CPUs come out of TSMC's constrained allocation. Intel owns fabs but has been working through yield problems and is moving capacity from PC chips to server chips [13].

A skeptic would call this a dinner-table story. On the type of evidence, that is fair. It rests on testimony from one room, one named CEO, one unnamed executive and one platform lead, and the newsletter does not include spot price data, a provider statement or a date for the change [2]. I think the direction holds anyway. The buyers describe from the demand side the same shortage Lesse explains from the factory side, and the constraints she lists are specific enough to check [9][11][13].

On timing, Lesse wrote that analysts expect CPU supply to gain headroom before memory does, but that relief is still multiple quarters away [14]. That puts the problem on a budget-cycle horizon. It is too long for a one-quarter workaround to cover, and short enough that a multi-year commitment could outlast the shortage. The trade-off is commitment against access. Reservations now have to be booked months ahead and can still be refused [4]. A team that books this quarter commits to a forecast of demand it has not yet seen.

What to watch

  • Published spot price histories or statements from AWS, GCP or Azure that confirm or contradict the dinner-table accounts of vanished CPU spot capacity.
  • Progress on Intel's yield problems and TSMC's CPU allocation, which would test the analysts' multiple-quarters timeline Lesse cited.
  • Whether reservation lead times stretch beyond months or refusals spread from large inference providers to mid-size accounts.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories