Skip to content

Build1 publisher3 min readPublished

Lambda ran 19 Blackwell nodes on the power budget provisioned for 16

The one measured number behind NVIDIA's tokens-per-megawatt pitch comes from Lambda's Blackwell cluster, where power reclaimed from static provisioning ran three extra nodes. Amazon's Annapurna Labs and d-Matrix got a line each.

The Engineer · Build desk

Illustration accompanying Lambda ran 19 Blackwell nodes on the power budget provisioned for 16

What happened

  • Ian Buck, NVIDIA's vice president of hyperscale and high-performance computing, spoke on AI factory efficiency at the AI Infra Summit in Santa Clara, which drew more than 8,000 attendees against 3,500 last year.
  • Lambda ran 19 nodes inside the power budget typically allocated to 16 full-power nodes, in what NVIDIA calls the first validation of its DSX MaxLPS power software on Blackwell servers.
  • Emerald AI and NVIDIA demonstrated automated load reduction at Silicon Valley Power, answering hundreds of demand signals while protecting AI workload performance.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision Grid-responsive power moves scheduling policy into the buyer's own priority tiers: the operator decides which internal customers feel a demand-response event.
  • exposure Any job an operator ranks below critical has its completion time partly set by the utility's signals, so batch and low-priority service commitments need rewriting before a flexible-load contract is signed.
  • capability NVIDIA lists NVLink as the scale-up interconnect in its own stack, so a Raptor part attaching through NVLink Fusion would make that fabric a socket for a competitor's inference silicon.
  • precedent A hyperscaler with its own accelerator team now co-develops NVIDIA's memory technology. That co-development sets the expectation that the largest buyers can specify custom HBM variants.

Static power provisioning hands every node a worst-case allowance and then leaves it there. NVIDIA's DSX MaxLPS monitors consumption across GPUs and racks, shifts available power to where it is needed most, and reclaims what static allocation leaves unused [9]. Training and inference draw differently over time, so a factory running both has more slack to move than one running either on its own [10].

Nineteen nodes inside the budget for sixteen is 18.75% more nodes [19]. Cluster throughput rose 24% and performance per watt 23% [8][6], which puts total draw about 1% above the baseline [20]. Divide the throughput gain by the node gain and per-node throughput is up about 4% [21]. Power capping does not normally make a node faster, so either the sixteen-node baseline was not throughput-saturated, or the rounding between "about 4 million" and "5 million" tokens per second absorbs the difference [8].

For 23% to show up in someone else's hall, several conditions have to match. The binding constraint has to be the electrical feed, because reclaimed watts only buy nodes when you cannot add them another way. Per-node allowances have to sit meaningfully above measured draw. And the workload has to be mixed, since that is where MaxLPS finds the imbalance to exploit [10].

NVIDIA's own figure is larger than the measured one. The company says MaxLPS can deliver up to 1.4x more tokens per megawatt through factory-wide power optimization [11], about 14% above Lambda's 1.23x [23]. The difference is scope: one number is the whole factory, the other is one cluster [11][6]. The Vera Rubin NVL72 claim in the same post breaks off mid-sentence at "up to 40% more GPU" [12].

The post also calls the summit an event "that has morphed into a Coachella of infrastructure tech" [3]. I have no way to validate that claim.

DSX Flex points the same control loop at the meter. It takes in load-shedding requests, demand-response events and pricing signals, then acts inside a hierarchy defined ahead of time: the most critical jobs keep going, everything else pauses temporarily and then resumes [16]. Emerald AI plans to run its Conductor grid-responsive power management software on Flex [15], and Silicon Valley Power operates the flexible-load interconnection program the demonstration ran against [14].

The two items with the longest reach are a sentence each. Amazon's Annapurna Labs is working with NVIDIA on NVHBM custom high-bandwidth memory [4], and d-Matrix is integrating the NVLink Fusion platform with its Raptor XPUs [5]. NVIDIA did not publish integration detail for either.

NVIDIA says the metric for AI infrastructure is shifting fast from peak performance to validated agentic tokens per megawatt [13]. Lambda's 23% is the one measured number in the slate [6]. Pinterest is described running the Blackwell platform and Dynamo inference software to bring conversational AI to visual discovery, with no figures attached [18].

What to watch

  • A second DSX MaxLPS validation on a single-profile inference fleet would show whether the 23% depends on mixing training and inference.
  • Whether NVHBM parts end up in Amazon's own accelerators, in NVIDIA's, or in both.
  • Whether Emerald AI's Conductor on DSX Flex converts the Silicon Valley Power demonstration into a paid grid-services arrangement.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories