Build1 publisherNot yet confirmed elsewhere3 min readPublished
Researchers aim AI data-center power cuts at server software
Researchers say software is the easiest fix for AI data-center power limits, citing a training optimiser that cut energy up to 30% on unchanged hardware. With average PUE flat for six years, the remaining gains are in the servers, which draw about 60% of a facility's power.
The Engineer · Build desk

What happened
- Nvidia says Blackwell power profiles save up to 15% of energy at 97% or more of performance, letting power-constrained sites run more GPUs for up to 13% more throughput.
- Cooling ranges from about 7% of demand at an efficient hyperscale site to more than 30% at a less efficient enterprise facility.
- ETH Zurich's Sophie Hall studies delaying or relocating batch jobs that are not time-sensitive, using day-ahead planning and real-time scheduling against grid signals.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint An operator who manages to PUE alone cannot see servers doing useless work, so a facility can meet its efficiency target while its biggest load runs wasted jobs.
- decision At an efficient hyperscale site a 15% server-side saving is worth more facility power than deleting cooling outright, so the next efficiency budget there goes further on software than on the plant.
- capability A power-capped site can add GPUs under the same feed; Nvidia's own inputs bound the gain at about 14% more throughput, close to the 13% it claims.
- constraint Perseus only saves where a training job is uneven and the FP8 figure covers one model, so neither number transfers to another workload without measuring it first.
PUE scores the building around the servers. It shows whether cooling and power systems are wasteful. It does not account for whether the software on those servers is doing useful work [4]. A server running dead code at full power leaves the ratio where it was [4].
At an efficient hyperscale site, cooling is about 7% of demand [6]. Nvidia says its Blackwell power profiles save up to 15% of energy [14]. Servers draw around 60% of demand [5]. Apply a 15% cut across that whole share and it comes to about 9% of facility power [18]. That is more than the same site would save by removing cooling entirely [19]. The comparison assumes the saving reaches every server. The profiles themselves tune GPUs: compute and memory frequencies, power limits, NVLink states and cache settings [13].
Nvidia's numbers are consistent with each other. Under a fixed power cap, 15% less energy per GPU leaves room for about 1.18 times as many GPUs. At 97% performance each, that is about 14% more throughput [17]. That only holds where a site is short of power and still has space and GPUs to add. The workload also has to keep 97% of its speed at the tuned settings. Nvidia claims up to 13% [14], just under that bound.
Perseus, Chung's training optimiser, finds the parts of a large-model training job that have less work to do. It slows them so they finish alongside the busier parts [11]. A part that finishes early only waits, so slowing it costs no throughput. The saving depends on imbalance, since a job whose parts carry equal work has nothing to slow. Perseus cut training energy by up to 30% this way without changing the hardware [12].
The FP8 figure needs the same caution. A slightly lower-precision format has to hold up on another model and prompt mix before that saving carries over. ML.Energy measured it on one model, Alibaba's Qwen 3 235B A22B Thinking, and only on problem-solving tasks [10].
Chung is a PhD candidate at the University of Michigan and a researcher with the ML.Energy initiative [8]. He treats computing as a stack: hardware at the bottom, then systems software, algorithms and applications [9]. Chips are slow and costly to replace. The three upper layers are easier to change, and gains there compound [9]. "The answer to what seems to be the core question - 'Can software or algorithms meaningfully help with power and energy?' - is yes, 100%," Chung said [7].
The other measures in the report are routine operations work. They are right-sized cloud instances, cached repeat prompts, small models for routine requests, better batching and compilers, and shorter prompts with capped outputs [16]. Chetan Visrolia of SHI has an audit that needs no tooling. He said the most obvious waste is often legacy equipment supporting old code [15]. "It is always the lowest-hanging fruit for quick savings," he said. "It can easily be found in any data center, as it is the loudest rack on the floor." [15]
Sophie Hall, a doctoral student at ETH Zurich, has studied workload shifting across Google's data-center fleet. She argues that consumption is not necessarily the main problem [1]. "It's more like: when do they use it, where do they use it, and how is it interacting with the grid?" she said [1].
Tom's Hardware did not put a price on any of these measures. The case that software is the cheaper route rests on the hardware half of Chung's stack argument [9].
What to watch
- Independent measurements of Nvidia's Blackwell power profiles on workloads Nvidia did not choose, to test the 97% performance figure.
- ML.Energy FP8 results on models beyond Qwen 3 235B A22B Thinking and on task types other than problem-solving.
- Whether the Uptime Institute's next survey reports any measure of useful work per watt alongside PUE.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+20
- Incentives50
- Confidence40
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
Sophie Hall, a doctoral student at ETH Zurich's Automatic Control Laboratory who has studied workload shifting across Google's global data-center fleet, argues electricity consumption is not necessarily the main problem: "It's more like: when do they use it, where do they use it, and how is it interacting with the grid?" she said.
ReportedSupportedSource: Sophie Hall, interview with Tom's Hardware Premium2 sources— create a free account to open themView cited source - [2]
Batch jobs that are not time-sensitive could be delayed until local demand falls or routed to a region with spare capacity and lower-carbon electricity; Hall's research uses day-ahead planning and real-time scheduling to respond to grid signals while preserving performance guarantees.
ReportedSupportedSource: Tom's Hardware, describing Sophie Hall's research2 sources— create a free account to open themView cited source - [3]
Average power usage effectiveness (PUE) has barely changed for six successive years, according to the Uptime Institute's 2025 survey.
- [4]
PUE can show whether cooling and power systems are wasteful but does not account for whether the software running on the servers is doing useful work.
- [5]
Servers account for around 60% of electricity demand in a modern data center.
- [6]
Cooling ranges from about 7% of demand in an efficient hyperscale site to more than 30% in a less-efficient enterprise facility.
- [7]
"The answer to what seems to be the core question - 'Can software or algorithms meaningfully help with power and energy?' - is yes, 100%," said Jae-Won Chung.
- [8]
Jae-Won Chung is a PhD candidate in computer science and engineering at the University of Michigan and a researcher with the ML.Energy initiative.
- [9]
Chung frames computing as a stack with hardware at the bottom, then systems software, algorithms and applications; hardware is hardest because chips are slow and costly to replace, while the other three layers are easier to change and gains there can compound.
- [10]
ML.Energy's tests of Alibaba's Qwen 3 235B A22B Thinking model found that running inference in FP8, a slightly lower-precision format, consumed a third less energy than bfloat16 versions on problem-solving tasks.
- [11]
Chung's Perseus training optimiser identifies parts of a large-model training job that have less work to do and slows them so they finish alongside busier parts.
- [12]
Perseus cut training energy by up to 30% without reducing throughput or changing the hardware.
- [13]
Nvidia's Blackwell power profiles fine-tune GPU compute and memory frequencies, power limits, NVLink states and cache settings to fit the workload.
- [14]
Nvidia believes Blackwell power profiles can save up to 15% of energy while retaining 97% or more of overall performance, allowing power-constrained facilities to run more GPUs and raise throughput by as much as 13%.
- [15]
Chetan Visrolia, business development manager for data center infrastructure at SHI, said the most obvious waste he encounters is often legacy equipment supporting old code: "It is always the lowest-hanging fruit for quick savings," he said. "It can easily be found in any data center, as it is the loudest rack on the floor."
ReportedSupportedSource: Chetan Visrolia, SHI, interview with Tom's Hardware PremiumView cited source - [16]
Other software measures include right-sizing cloud instances, caching repeatedly used prompts, using small models for routine requests, better batching and compilers to raise accelerator utilisation, and compressing prompts and limiting outputs to cut tokens processed.
- [17]
Under a fixed power cap, a 15% per-GPU energy saving allows about 1.18 times as many GPUs; at 97% performance each, throughput rises about 14%, just above Nvidia's claimed 13%.
- [18]
A 15% cut applied across a server share of about 60% of facility demand is about 9% of facility power.
- [19]
At an efficient hyperscale site, a 9% facility-level saving from servers exceeds the roughly 7% that cooling draws in total.
Sources
1 independent publisher whose own reporting we read for this story.
- tomshardware.comSoftware could be the easiest fix for hyperscalers' AI power squeeze, researchers say
1 article · October 8, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.