Build1 distinct publisher2 min readPublished
The zettaFLOPS figure is an estimate. The measured one, sitting in a GB300 example, is a 25 percent density gain, and it is paid for out of the safety margin.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Divide the two headline numbers and you get the figure an operator will actually plan against: 2.5 kW of site power for every Rubin GPU installed [1], or 2.5 MW per SuperPOD [2]. That is measured at the meter, so it has to cover cooling, networking, conversion losses and everything else in the shell, not just the accelerator. Run the same division on the performance claim and the site delivers roughly 20 TFLOPS per watt for NVFP4 inference and 14 for training [3]. Those are the numbers to hold Nvidia to, and Tom's Hardware has already said the FLOPS half is estimated rather than measured, with real workloads landing significantly lower [3].
The measured part of the pitch is smaller and more useful. In Nvidia's GB300 example, provisioning statically against a 135 kW peak fills a 540 kW budget with four racks and leaves as much as 170 kW unspent [4], which is 31 percent of the allocation sitting behind a breaker doing nothing [5]. Reclaim it and a fifth rack fits inside the same budget [5]. That is 25 percent more installed compute per contracted megawatt [6], at an average allocation of 108 kW per rack rather than 135 [7].
Where that 25 percent comes from decides whether it travels. The control loop harvests the gap between what a group of racks could draw and what it does draw, watching power at chip, rack and group-of-racks level and moving slack to whichever system needs it [6]. Nvidia's own description of the problem is a load differential between racks, one busy and one not [7]. That differential is a property of the workload mix, not of the silicon. A hall running one synchronous training job across every rack has less to redistribute than a mixed inference fleet with idle capacity, and the rack-level profiles Nvidia ships for inference, training and memory-bound work [8] are an admission that the answer changes by job type.
Then there is the reserve. The five-rack configuration works partly because much less power is held back for peaks [5]. Static provisioning was wasteful, and it also failed safe: that stranded 170 kW was, incidentally, insurance. Dynamic allocation converts the insurance into inventory and moves the guard band into software. That is a defensible trade for an operator who can measure it. Note also what DSX stands for, which is Land, Power, Shell [1], and what that implies about where Nvidia now expects to be consulted.
Ranked by verification strength, evidence, and original report placement.
Tom's Hardware, doing its own back-of-the-napkin math from publicly available Rubin specs, said it felt safe assuming those performance figures are estimated rather than measured, and that the maximum achievable FLOPS from real-life workloads is likely to be significantly lower.
At its Hot Chips presentation on the Rubin GPU, Nvidia said that all of Vera Rubin's power management technologies combined with its DSX MaxLPS (Land, Power, Shell) suite of design and site-level dynamic power management resources will let operators provision 40,000 Rubin GPUs, about 40 Rubin DGX SuperPODs, within an example fixed facility power budget of 100MW.
Nvidia expects that hardware to deliver up to 2 zettaFLOPS for NVFP4 inference and up to 1.4 zettaFLOPS for NVFP4 training.
In Nvidia's measured example using current GB300 racks, an operator statically provisioning for peaks could install four 135kW systems within a 540kW power budget, but in practice as much as 170kW of that budget might sit unused because of differences in rack utilisation.
Nvidia's example shows five rack-scale systems installed within the same 540kW budget, and across those five systems the amount of reserve power allocated for peaks can be much lower.
DSX MaxLPS applies a dynamic scheme continuously aware of power usage at chip level, rack level and groups-of-racks level; Nvidia's Dynamic Power Software control loop finds unused power and redistributes it to systems where it is most needed at that moment.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Vendor figures, one outlet, partly self-flagged as estimates
Everything traces to a single trade-press write-up of an Nvidia conference session plus a companion Nvidia blog post. The GB300 kilowatt example is described as measured and is internally consistent enough to support arithmetic (31 percent stranded, 25 percent more racks, 108kW average per rack), which raises the floor. The forward-looking Rubin numbers are the weak part: the publisher itself concludes they are estimated rather than measured, and no independent measurement, third-party audit or second outlet corroborates any figure.
Announcement stage, no disclosed deployments
The supplied material documents a conference presentation and a vendor-run prior-generation characterization on GB300 hardware, and nothing further: no named operator, colo or cloud tenant, no shipping schedule, no pricing or licensing terms, and no measured Rubin deployment. Adoption is therefore real but confined to vendor demonstration and toolkit availability messaging.
Headline forecast outruns the measured example
The distributable number - 2 zettaFLOPS of inference in 100MW - is an unmeasured vendor estimate for hardware with no disclosed measured power or performance-per-watt data, while the substantiated result is narrower: a 25 percent increase in racks per contracted megawatt in a GB300 example, achieved by cutting the average per-rack envelope 20 percent below the static peak allocation. That is a real and useful gain, but it is funded from the safety margin and the source does not examine the reliability cost. The gap is moderated, not erased, by the publisher labelling the estimates as estimates in the same article.
Vendor-authored numbers at a launch moment
Every quantitative input originates with Nvidia - a conference session on its own next-generation GPU, a companion Nvidia blog post, and Nvidia's own GB300 characterization - and each supports a commercial conclusion Nvidia benefits from: that more of its accelerators fit inside a fixed power contract and that its DSX toolkit is the way to get there. The report also states Nvidia did not release measured Rubin power or performance-per-watt data, so the selection of what was disclosed is vendor-controlled. The publisher's independent caveat and its own napkin math offset this only partly.
Specific and self-caveated, but single-sourced
Confidence is moderate: the reporting is concrete, internally consistent and unusually candid about the estimate/measurement boundary, which makes the GB300 arithmetic dependable as a description of Nvidia's claim. But there is one publisher, one underlying vendor session, no independent verification, no adoption data, and the source body is truncated mid-argument on lifecycle re-provisioning, so several dimensions rest on a narrow base.
build
Nvidia's 88-core Vera bets agentic serving runs out of bandwidth before it runs out of cores1 distinct publisher
build
A benchmark that replays real agent sessions gives back less of the generational win2 distinct publishers
build
NVIDIA's new capacity math: 60 of every 100 megawatts actually reaches the AI load1 distinct publisher
build
Nvidia's Groq-derived LPX rack posts 3,431 tokens/sec on 128GB of SRAM1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.