Skip to content

Build1 publisherNot yet confirmed elsewhere3 min readPublished

Researchers aim AI data-center power cuts at server software

Researchers say software is the easiest fix for AI data-center power limits, citing a training optimiser that cut energy up to 30% on unchanged hardware. With average PUE flat for six years, the remaining gains are in the servers, which draw about 60% of a facility's power.

The Engineer · Build desk

How we use AISend a correction

Illustration accompanying Researchers aim AI data-center power cuts at server software
Generated illustration

What happened

  • Nvidia says Blackwell power profiles save up to 15% of energy at 97% or more of performance, letting power-constrained sites run more GPUs for up to 13% more throughput.
  • Cooling ranges from about 7% of demand at an efficient hyperscale site to more than 30% at a less efficient enterprise facility.
  • ETH Zurich's Sophie Hall studies delaying or relocating batch jobs that are not time-sensitive, using day-ahead planning and real-time scheduling against grid signals.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint An operator who manages to PUE alone cannot see servers doing useless work, so a facility can meet its efficiency target while its biggest load runs wasted jobs.
  • decision At an efficient hyperscale site a 15% server-side saving is worth more facility power than deleting cooling outright, so the next efficiency budget there goes further on software than on the plant.
  • capability A power-capped site can add GPUs under the same feed; Nvidia's own inputs bound the gain at about 14% more throughput, close to the 13% it claims.
  • constraint Perseus only saves where a training job is uneven and the FP8 figure covers one model, so neither number transfers to another workload without measuring it first.

PUE scores the building around the servers. It shows whether cooling and power systems are wasteful. It does not account for whether the software on those servers is doing useful work [4]. A server running dead code at full power leaves the ratio where it was [4].

At an efficient hyperscale site, cooling is about 7% of demand [6]. Nvidia says its Blackwell power profiles save up to 15% of energy [14]. Servers draw around 60% of demand [5]. Apply a 15% cut across that whole share and it comes to about 9% of facility power [18]. That is more than the same site would save by removing cooling entirely [19]. The comparison assumes the saving reaches every server. The profiles themselves tune GPUs: compute and memory frequencies, power limits, NVLink states and cache settings [13].

Nvidia's numbers are consistent with each other. Under a fixed power cap, 15% less energy per GPU leaves room for about 1.18 times as many GPUs. At 97% performance each, that is about 14% more throughput [17]. That only holds where a site is short of power and still has space and GPUs to add. The workload also has to keep 97% of its speed at the tuned settings. Nvidia claims up to 13% [14], just under that bound.

Perseus, Chung's training optimiser, finds the parts of a large-model training job that have less work to do. It slows them so they finish alongside the busier parts [11]. A part that finishes early only waits, so slowing it costs no throughput. The saving depends on imbalance, since a job whose parts carry equal work has nothing to slow. Perseus cut training energy by up to 30% this way without changing the hardware [12].

The FP8 figure needs the same caution. A slightly lower-precision format has to hold up on another model and prompt mix before that saving carries over. ML.Energy measured it on one model, Alibaba's Qwen 3 235B A22B Thinking, and only on problem-solving tasks [10].

Chung is a PhD candidate at the University of Michigan and a researcher with the ML.Energy initiative [8]. He treats computing as a stack: hardware at the bottom, then systems software, algorithms and applications [9]. Chips are slow and costly to replace. The three upper layers are easier to change, and gains there compound [9]. "The answer to what seems to be the core question - 'Can software or algorithms meaningfully help with power and energy?' - is yes, 100%," Chung said [7].

The other measures in the report are routine operations work. They are right-sized cloud instances, cached repeat prompts, small models for routine requests, better batching and compilers, and shorter prompts with capped outputs [16]. Chetan Visrolia of SHI has an audit that needs no tooling. He said the most obvious waste is often legacy equipment supporting old code [15]. "It is always the lowest-hanging fruit for quick savings," he said. "It can easily be found in any data center, as it is the loudest rack on the floor." [15]

Sophie Hall, a doctoral student at ETH Zurich, has studied workload shifting across Google's data-center fleet. She argues that consumption is not necessarily the main problem [1]. "It's more like: when do they use it, where do they use it, and how is it interacting with the grid?" she said [1].

Tom's Hardware did not put a price on any of these measures. The case that software is the cheaper route rests on the hardware half of Chung's stack argument [9].

What to watch

  • Independent measurements of Nvidia's Blackwell power profiles on workloads Nvidia did not choose, to test the 97% performance figure.
  • ML.Energy FP8 results on models beyond Qwen 3 235B A22B Thinking and on task types other than problem-solving.
  • Whether the Uptime Institute's next survey reports any measure of useful work per watt alongside PUE.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence45
Adoption
Insufficient
Hype gap+20
Incentives50
Confidence40
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    Sophie Hall, a doctoral student at ETH Zurich's Automatic Control Laboratory who has studied workload shifting across Google's global data-center fleet, argues electricity consumption is not necessarily the main problem: "It's more like: when do they use it, where do they use it, and how is it interacting with the grid?" she said.

    ReportedSupportedSource: Sophie Hall, interview with Tom's Hardware Premium2 sources— create a free account to open themView cited source
  2. [2]

    Batch jobs that are not time-sensitive could be delayed until local demand falls or routed to a region with spare capacity and lower-carbon electricity; Hall's research uses day-ahead planning and real-time scheduling to respond to grid signals while preserving performance guarantees.

    ReportedSupportedSource: Tom's Hardware, describing Sophie Hall's research2 sources— create a free account to open themView cited source
  3. [3]

    Average power usage effectiveness (PUE) has barely changed for six successive years, according to the Uptime Institute's 2025 survey.

    ReportedSupportedSource: Tom's Hardware, citing Uptime Institute 2025 surveyView cited source

Sources

1 independent publisher whose own reporting we read for this story.

  1. tomshardware.com

    1 article · October 8, 2026

    Software could be the easiest fix for hyperscalers' AI power squeeze, researchers say

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Topics

  • Data-center energy efficiencyFollow
  • AI inference efficiencyFollow
  • Carbon-aware workload schedulingFollow
Loading related stories