Leadership1 publisher3 min readPublished
An HPE practice leader wants business leaders to meter AI capacity in dollars per delivered FLOP
Vinod Bijlani argues in Forbes that a token price measures only what an application consumed, while the idle accelerators and rising power draw behind it never reach the invoice a CIO reviews. He works for HPE.
The Board Room · Leadership desk

What happened
- Vinod Bijlani, an AI practice leader at Hewlett Packard Enterprise, argued in Forbes that price per million tokens measures the price of consuming AI while leaving the productivity of the infrastructure producing it unmeasured.
- The International Energy Agency projects global data center electricity consumption roughly doubling from 485 terawatt-hours in 2025 to around 950 by 2030, with AI-focused sites roughly tripling.
- Bijlani puts idle accelerators, memory constraints, network contention and energy inefficiency under $/FLOP, separate from token consumption and from the economics of usable data.
Compiled by The Board RoomSomething wrong?How this is made
Why it matters
- constraint If power, cooling and floor space are the binding limits, the purchase question stops being how many accelerators a budget buys and becomes how much reliable computation fits inside the site.
- decision Three separate meters put a name on each overrun: the platform team answers for idle silicon and contention, the application team for context growth and extra model calls.
- cost Underutilization is charged to whoever holds the hardware, so a firm that rents by the token pays for it inside the unit price, where there is no line item to interrogate.
- contradiction Efficiency and demand are compounding at once, so a board that watches unit costs fall learns nothing about whether its own usage is under control.
Put the two Epoch AI series on one axis and the efficiency case is thinner than it first reads. Chip spending buys about 49% more performance each year [4], and Epoch's 2024 pretraining work found the compute needed to reach a given performance level halved roughly every eight months [5], which works out to about 2.8 times a year [1]. Compound the two and capability per dollar improves roughly 4.2 times a year [2]. Frontier training compute has grown about five times a year since 2020 [6]. Five divided by 4.2 leaves spending rising about 19% a year [3]. Epoch's demand figure is a frontier training series, so that 19% belongs to a lab, not to an enterprise inference budget. The direction is what Bijlani claims for the enterprise: lower unit costs make longer contexts, more model calls, richer modalities and larger agent loops economically possible [14].
Whose problem $/FLOP is depends on who holds the hardware. Buy inference from a provider by the million tokens and the underutilized accelerator sits on the provider's books, priced into the token whether the buyer can see it or not. Own or reserve the capacity and the same idle silicon lands in your depreciation and your power bill. The column addresses CIOs and AI leaders as one audience, though only some of them carry that exposure. Bijlani is an AI practice leader at Hewlett Packard Enterprise [1], and a metric that scores owned, well-loaded infrastructure is a comfortable one to recommend from that seat.
The International Energy Agency projects global data center electricity consumption roughly doubling from 485 terawatt-hours in 2025 to around 950 by 2030 [7], about 14% a year compounded [4], with AI-focused sites growing faster and roughly tripling over the same period [8], close to 25% a year [5]. That forecast is the IEA's, independent of Bijlani's employer. Bijlani wrote that the more pressing question is "How much reliable computation can we deliver within the power, cooling and space actually available?" [10]
He splits the ledger three ways: poor data quality to datanomics, excessive prompts and context growth and model calls to tokenomics, and idle accelerators, memory constraints, network contention and energy inefficiency to $/FLOP [11]. That governance point survives the vendor interest. Collapse the three into one number and the overrun floats free of the team that caused it. "Falling $/FLOP doesn't guarantee a falling AI bill, and it can't tell leaders whether demand is being controlled," he wrote [9].
He argues $/FLOP should not be used as an isolated purchasing score, and that the useful enterprise number is effective cost per delivered FLOP for a defined workload and service level [12]. That is harder to put in a purchase order than a chip spec. Published peak performance overstates what gets delivered: architecture, numerical precision, batch size, memory bandwidth, communication overhead and utilization all decide how much installed capacity becomes productive work [13].
What to watch
- Whether Epoch AI's 49%-a-year price-performance series holds through the next accelerator generation, since the efficiency half of the argument rests on it.
- Whether the IEA revises the 950 terawatt-hour 2030 projection, and whether AI-focused sites track the tripling.
- Whether infrastructure vendors and cloud providers start quoting effective cost per delivered FLOP with a service level attached to it.