Build1 distinct publisher3 min readUpdated
A power-budget waterfall in NVIDIA's DSX MaxLPS post puts facility overhead, rack losses and restart inefficiency at 40% of grid input. Planners should price the remainder, not the racks.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
NVIDIA has published a design pitch for what it calls DSX MaxLPS, and the framing matters more than the product: the company argues the capacity question is no longer how many GPUs fit in a data center but how much AI output each available megawatt delivers [1]. Buried in the post is the number operators should be modelling against, which is that in one representative power-budget view examined by NVIDIA, roughly 60% of delivered site power is allocated to compute for AI output [4].
The arithmetic is laid out explicitly for a 100 MW site. Of 100 MW of grid input, 20 MW goes to facility overhead, 10 MW to rack losses, and 10 MW is unavailable to the AI load because of operational inefficiency during failures, restarts and checkpointing, leaving 60 MW [5]. That is 40 MW of every 100 MW that never becomes tokens [6]. Read as a procurement rule, adding 10 MW of AI load at those same ratios means buying about 16.7 MW of grid input [7]. Facility overhead alone is twice the size of either the rack-loss or the restart-inefficiency slice [18].
Power distribution, cooling, networking, storage, backup and facility systems all take their share before electricity reaches a GPU [3]. On top of that, NVIDIA singles out static rack provisioning: the practice of allocating each rack its maximum possible draw to cover worst-case peaks, with further capacity reserved for failures, operational flexibility and expansion [8]. The consequence is that each rack becomes an isolated power island, and a rack sitting on unused reserve cannot lend it to a neighbour that would convert it into output [9]. That is not a rounding error, because training, post-training and inference cycle through compute bursts, memory-bound execution, synchronization, checkpointing, prefill, decode, idle gaps and network-bound communication, each drawing power differently, so a rack provisioned for its peak runs below that level for meaningful periods [10].
MaxLPS, which NVIDIA expands as Maximum Land Power Shell, is described as a suite of chip, thermal, system and software technologies aimed at maximising throughput inside a fixed power budget [11]. It works on three layers: dynamic power allocation, software performance-per-watt techniques, and 45C warm-water liquid cooling with site design changes intended to improve PUE [12]. The allocation piece is Dynamic Power Software, currently in Developer Preview, which models the site topology from utility down to racks, nodes and GPUs while operators set resource groups, budgets and policies [13]. It compares allocated against actual consumption and hands headroom from underused GPUs and racks to others in the same managed group, leaving the site envelope unchanged [14], running a continuous loop of telemetry collection, reallocation within policy, and validation against the approved group budget [15].
Note where that acts. NVIDIA presents the stranded rack headroom as a separate problem from the facility overhead, rack losses and restart inefficiency that shrink the AI load in the first place [16], which means dynamic allocation is redistributing inside the 60 MW rather than clawing back the 40 MW [19]. Only the thermal layer is described as attacking cooling overhead directly [12]. And the post does not publish a recovered-megawatt or percentage throughput figure for any of the three layers [20].
For inference specifically, NVIDIA nominates application-level performance per watt as the efficiency metric [2]. That is the number to demand in any vendor comparison, because it is the only one that survives contact with a fixed shell.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
In one representative power-budget view examined by NVIDIA, about 60% of delivered site power is allocated to compute for AI output.
NVIDIA's power-budget waterfall for a 100 MW AI factory allocates 20 MW of the 100 MW grid input to facility overhead, 10 MW to rack losses, and 10 MW unavailable for AI load because of operational inefficiency during failures, restarts and checkpointing, leaving 60 MW available for AI load; each deduction is expressed as a share of the original grid input.
NVIDIA describes static rack provisioning as an outdated approach that allocates the maximum power draw per rack to meet worst-case peak demand even though real workloads have different power needs and may leave some of that maximum unused, with operators reserving additional capacity for failures, operational flexibility and expansion.
Traditional power planning treats each rack as an isolated power island, and an isolated rack provisioned with excess power cannot lend that unused power to a neighbour that could turn it into tokens.
DPS continuously compares allocated power against actual consumption and, when GPUs or racks operate below their reserved level, makes that headroom available to others in the same managed group, leaving the site power envelope unchanged.
DPS runs a continuous control loop that collects GPU-, rack- and group-level power telemetry, identifies unused power capacity, reallocates power within policy, and validates compliance with the approved group power budget.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Single vendor-authored source with illustrative figures
Every claim traces to one NVIDIA developer blog post. The power-budget waterfall is explicitly a 'representative' view and the efficiency comparison is an illustration, not measured production data; there is no independent measurement, third-party benchmark or customer telemetry in the cluster. The internal descriptions are detailed and self-consistent, which supports what NVIDIA asserts, but not the real-world magnitude of the gains.
Developer Preview only, no disclosed deployments
The two software components that carry the argument — DPS and DSX Exchange — are stated to be in Developer Preview, and the post names no operator, site or production deployment using MaxLPS dynamic power allocation. Adoption evidence is therefore limited to vendor announcement and preview availability.
Framing outruns the published numbers
The post promises to 'maximize AI factory throughput within a fixed power budget' and casts static provisioning as outdated, yet the mechanism it details redistributes headroom inside the 60 MW already allocated to AI load rather than recovering the 40 MW deducted upstream, and no site-scale gain is quantified for any of the three layers. The scale-invariant read of the single rack-level example (170 kW of a 540 kW budget) is left to the reader while the headline framing is site-level, and the software is pre-GA. The underlying capacity math itself is sober and useful, which keeps the gap moderate rather than extreme.
Vendor selling the remedy defines the constraint
The sole source is NVIDIA's own developer blog. NVIDIA both frames the bottleneck (megawatts, not GPU count; static rack provisioning as outdated) and sells the described answer — the DSX MaxLPS suite, DPS software, 45C liquid-cooled site design and workload power profiles. The efficiency numbers, the representative waterfall and the static-versus-dynamic comparison are all vendor-constructed, with no disclosed independent validation.
High confidence in what was said, low in what it yields
Confidence is high that NVIDIA published these figures and mechanisms, since the primary vendor document states them plainly and quantitatively. Confidence is low on the operational outcome: one publisher, pre-GA software, representative rather than measured numbers, and one internal inconsistency about whether quantified recovery exists (site-scale absent, rack-scale present).
build
China's accelerator swap makes Cambricon supply, not export policy, your ship-date risk1 distinct publisher
product
Nebius funds $4.5bn of AI capacity on terms that pay lenders mostly in stock2 distinct publishers
build
NVIDIA put a number on agent skills: 300+ verified, two harnesses, baselines under 50/1001 distinct publisher
invest
The chips never move: Washington's fix for the Southeast Asia compute loophole1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 21, 2026