Skip to content

Product1 publisher2 min readPublished

Omdia's analysts put the bottleneck diagnosis ahead of the next rack purchase

In a TechTarget feature, two Omdia analysts argue that installed infrastructure and usable capacity are different things, and they name the power, storage, network and cooling dependencies that separate the two.

The Product Desk · Product desk

Photograph accompanying Omdia's analysts put the bottleneck diagnosis ahead of the next rack purchase
Photo: datacenterfrontier.com

What happened

  • A TechTarget feature argues that the clash between aggressive new data center investment and underused existing sites comes from treating installed infrastructure and usable capacity as the same thing.
  • Its central example is a facility with open rack space that lacks the power the target workload needs.
  • Bruce Bateman, Omdia's chief analyst for semiconductors, cautioned that upgrading compute alone does not drive utilization and that not every legacy site should become a high-density AI facility.
  • Omdia cloud and data center analyst Siraj Aziz said dynamic and agentic AI workloads complicate capacity planning because their processing and hardware needs vary over time, with visibility fragmented across systems.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • cost The capital for UPS systems, electrical distribution and cooling is already spent, and the maintenance, staffing, connectivity and monitoring bills arrive whether or not the racks carry work.
  • constraint One binding resource pulls down the productive value of the assets behind it, so a purchase order for compute cannot buy back a facility whose real limit is distribution or cooling.
  • decision Modernization gets priced as a total retrofit against a named target workload. The electrical and cooling work lands on the same page as the server quote.

A colocation provider can walk a customer past rows of empty racks and still be unable to sell them, because the power density or configuration that customer needs is not there [15]. Those racks exist physically, and they never turn into revenue-generating inventory [15]. An enterprise hits the same wall in its own room when its existing systems cannot meet the workload's requirements [16].

The number that misleads is utilization. Low CPU or GPU utilization does not mean there is compute to spare, since the processor may be waiting on memory, storage, networking or another stage of the workload pipeline [3]. Capacity has to be evaluated across the system [10]. Some of the slack is also on purpose: operators provision above normal demand to absorb peaks and protect application performance, and application architecture and orchestration limit how efficiently what is available gets allocated [11].

Compute is the part of this with a quote attached. Bateman said that adding newer compute also requires more storage, higher-capacity routers and internet connectivity, and that upgrading for higher-density AI infrastructure needs changes to electrical distribution and cooling systems; retrofitting liquid cooling can introduce new plumbing, permitting and facility requirements [4]. Training clusters ask for all of it at once, coupling accelerators with high-bandwidth networking, substantial storage throughput, high rack power density and advanced cooling [13]. Inference is less predictable, its requirements moving with the model, the concurrency and the latency target, so a floor that runs traditional enterprise applications cannot be assumed to hold equivalent capacity for every AI workload [14].

Aziz put the question from the business end. Productive capacity depends on whether systems and workloads can use the installed resources to meet performance and business requirements, and, he said, the more important question is why an organization cannot put more workloads on capacity that already exists, and what that mismatch costs the business [18].

The feature carries no numbers [20]. There is no sample of facilities, no measure of how much installed capacity sits stranded and no figure for what a single bottleneck costs, so what is on offer is an order of operations, and the assertion that many existing facilities have significant underused capacity is TechTarget's own [1][20].

That order of operations works without new data. An operator names the tenant and the workload, then identifies which resource binds first: power, cooling, network, storage or the application itself. Bateman said an operator should determine who will use the capacity and for what workload [8].

What to watch

  • An Omdia dataset naming which resource binds first across a measured sample of facilities.
  • Colocation providers reporting sellable power capacity separately from rack and floor space counts.
  • Liquid-cooling permitting timelines appearing as a named line in operators' retrofit cost disclosures.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories