Build1 publisher3 min readPublished
Nvidia's MaxLPS trial at Nscale ran 192 GPUs against a 140-GPU static baseline
Nvidia says DSX MaxLPS power sharing fits up to 40% more GPUs in the same approved power budget, citing an Nscale trial of 192 GPUs against 140. The gain holds only while GPUs rarely peak together, and each operator has to measure that on its own fleet before buying hardware against it.
The Engineer · Build desk

What happened
- Nvidia's static baseline ran two 52-GPU high-throughput Kimi K2.5 instances and one 36-GPU low-latency instance on 35 four-GPU nodes.
- The MaxLPS configuration used 48 four-GPU nodes and put the extra capacity into a third 52-GPU high-throughput instance, keeping the low-latency instance at 36 GPUs.
- Nscale ran the MaxLPS software and collected telemetry at its Verne campus site in Keflavik, Iceland, while Nvidia ran the workloads on GB300 NVL72 systems.
- The test served Kimi K2.5 in FP4 on Blackwell Ultra GPUs with Dynamo and TensorRT LLM, using 8K-token inputs and 1K-token outputs.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision Before adding nodes to a pool, an operator has to write the group limits, reserve requirements, priorities and emergency responses that per-node peak sizing never asked it to define.
- exposure The low-latency service competes for power with 156 high-throughput GPUs, so its latency in a budget squeeze depends on the priority the operator assigns it in policy.
- constraint Operators running mostly training would be extrapolating from an inference-only trial, so they need their own test before counting on extra GPUs in the same budget.
Static planning reserves enough power for every node to reach its specified peak at the same time [5]. The reservation is per node. Headroom that one node leaves unused cannot be lent to another, so the facility can sit below its limit while GPUs that would fit stay offline [6]. Nvidia's case rests on how rarely those peaks coincide. Inference alternates among prefill, decode, memory-bound work, network activity and idle intervals, and two instances of the same model draw different power as request shapes and concurrency change [7].
MaxLPS replaces the per-node reservation with a pooled one, and Dynamic Power Software is the control layer [15]. Nvidia describes the loop in five parts [8]:
1. Operators map participating nodes into a managed group with one aggregate power budget. 2. Telemetry is collected at GPU, node, rack and group level, often enough to see headroom and emerging power events. 3. Policy sets node limits, group limits, allocation priorities, reserve requirements and responses to maintenance or emergencies. 4. When some resources draw less than their allocation, the software adjusts GPU power limits so others can use the slack. 5. Enforcement compares measured power with the group budget and adjusts allocations as consumption approaches it.
I think the design is sound. The approved budget stays fixed and is checked against measured power; only the per-GPU limits move [8]. "This is coordinated allocation, not an increase in the site's power supply," Nvidia wrote [9].
The group budget is not the only ceiling. Utility service, substations, distribution equipment, racks, nodes and GPUs each impose a limit, and every managed boundary has to stay inside its approved figure [4]. The trial kept each distributed workload inside one of four racks in both configurations [13]. Nvidia said that controlled for cross-rack performance differences [13].
The larger configuration has 1.37 times the baseline's GPUs, or 37% more [1]. That is a little under the "up to 40%" in Nvidia's headline claim [1], which says the gain comes inside the same approved power budget [1]. At a fixed budget, the average allowance per GPU at full draw is 140/192 of the static reservation, about 73% [3]. If the budget was sized for the baseline's peak, all 192 GPUs cannot draw that peak at once [3]. When combined consumption approaches the budget, the enforcement step pulls allocations back [8].
The software moves power only when some resources draw less than their allocation [8]. A fleet whose nodes peak together gives it little to move; that fleet is also the one static planning already sizes correctly. For the 37% to carry to another site, its instances have to peak at different times, as this trial's mix of high-throughput and low-latency inference instances with distinct power and service profiles did [10].
The post says it covers the measured trade-offs between power and performance, the controls used to hold electrical limits, and a repeatable validation method operators can use before deploying at scale [3]. The copy of the post reviewed for this article breaks off mid-sentence in the methodology section, before any throughput, latency or power-draw figures [16].
What to watch
- The measured throughput and latency results from the Nvidia-Nscale evaluation, especially for the 36-GPU low-latency instance under MaxLPS.
- Any MaxLPS trial on training jobs, which cycle through compute, communication, synchronization and checkpointing and were not part of this test.
- Whether Nscale or another operator runs Nvidia's validation method on its own fleet and publishes per-GPU power and throughput figures.