Leadership1 distinct publisher3 min readUpdated
Estimates for the water behind one GPT prompt run from over 500ml down to 0.3ml, depending on where the boundary is drawn. Fleet totals and efficiency ratios are what a board can actually audit.
The Board Room · Leadership desk

Compiled by The Board RoomSomething wrong?How this is made
The 500ml estimate and the 0.3ml estimate are not two measurements of one thing. According to the Forbes account, the gap opens up depending on whether upstream cooling is counted, what the task actually is, and where the demand originates [4]. Each of those is a scoping decision taken before anything is measured, which is how two people working honestly end up roughly 1,667 times apart on the same prompt [12].
That makes the per-query figure unusable as a governance object. Put one in a sustainability report and you have published a number whose opposite is equally citable from the same accounting menu [3][4]. The Forbes piece reports that a fair number of experts have already written the exercise off as futile and would rather push on efficiency [5]. Fine, but efficiency needs a denominator someone owns.
The macro side is not automatically that denominator. Yankai Jiang's estimate of total data center consumption at around 4.5 terawatts, glossed as "7 New York Cities" of power, came from a talk at Planet Action [6], an event the article's author says he helps run [7]. It is stated as a power figure with no reporting period attached, and it is an estimate from a stage rather than a company filing, which is precisely the distinction Jordan Hale of Tech Journal draws when he argues aggregate disclosures are more reliable because companies report them directly [11][14]. The auditable quantity is not global consumption. It is what your named suppliers report, what share of it you contracted for, and whether the ratio of output to input moves year over year.
The design recommendations in the same piece are more useful to a leadership team than any per-prompt figure, because they are contract language rather than disclosure prose. Jiang's Arizona-versus-Finland contrast turns on siting and cooling method: the Finnish case uses cold climate and nearby seawater and runs jobs when renewable energy peaks, without fresh water [8]. He treats carbon-aware scheduling as a partial but worthwhile start [9]. He also argues against overprovisioning hardware, in his words, "You do not always need a Ferrari. Sometimes a reliable Toyota will do just fine" [10]. A region, a cooling type, a scheduling window and a machine class are all things a procurement team can specify and later verify.
There is a cost line hiding in his rhetoric too. Jiang notes that when data centers burn out, "we just retire the old ones and build new ones" [13]. Replacement cadence is capital expenditure, and capital expenditure is already reviewed by the people being asked to defend millilitres per prompt.
The per-query number will keep circulating because it is vivid and free to quote. It is also the only figure in this debate that cannot survive a second question about its boundary.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Published estimates run from over 500ml of water per GPT query down to 0.3ml, the low figure attributed to Sam Altman.
The variation in per-query estimates depends on factors including whether upstream cooling is accounted for, what the task is, and where the demand is coming from.
Jiang contrasted a hot, dry, water-stressed Arizona data center with a Finnish one that uses cold climate and nearby seawater for cooling without using fresh water, and that runs jobs when renewable energy peaks.
Jiang said carbon scheduling is only part of the equation but is a good start.
Jiang recommends combining hardware according to need, saying "You do not always need a Ferrari. Sometimes a reliable Toyota will do just fine."
The article's author states that he helps run Planet Action, the event where Jiang spoke.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Thin and single-sourced
One publisher, one item. The measurement critique is internally verifiable — the article states both endpoints of the per-query range and the ratio follows arithmetically — but every substantive quantitative claim rests on an unmethodologised conference talk or an unlinked second-hand quotation. No company disclosure, dataset, facility name, reporting period or counter-expert appears anywhere in the supplied material.
No adoption evidence supplied
The supplied material contains no release, deployment, benchmark, disclosure or usage event. The Arizona and Finland contrasts are personified illustrations ('Dave, who lives in Arizona'), not identified facilities, and no operator is named as having adopted seawater cooling, carbon-aware scheduling or heterogeneous hardware fleets. Inferring uptake from a design prescription would be guessing.
Prescription outruns the evidence
The cluster framing moves from a sound observation — a 1,667x spread is not a governable metric — to a governance prescription that fleet totals and efficiency ratios are what a board can audit. But the only fleet total on offer is a single conference estimate stated as raw power with no reporting period, which fails the company-reported reliability test the piece itself endorses. The headline confidence and the audit framing exceed what one uncorroborated contributor column supports, though the source is more hedged than the framing ('the numbers are slippery').
Disclosed but material author-source tie
The author states plainly that he helps run Planet Action, the event at which his principal expert source spoke — a disclosed but real incentive to amplify that speaker's framing and figures, and the piece contains no source obtained independently of that event apart from one quoted line from another outlet. Disclosure earns credit; the absence of any independent corroboration limits it.
Low
Confidence is limited by single-publisher sourcing, the absence of any adoption evidence, an unresolved unit/period ambiguity in the one aggregate figure, and a disclosed author-source relationship. What can be held with reasonable confidence is narrow: that the circulating per-query estimates diverge by roughly three orders of magnitude, and that the divergence tracks where the accounting boundary is drawn.
invest
The 81% Problem: AI's Star CEOs Are Polling Badly With The People They Need To Hire1 distinct publisher
product
Washington's secret AI test is coming for open weights, and release dates go with it2 distinct publishers
leadership
The same lies, every fire season: a coalition report makes the case for pre-planned crisis comms1 distinct publisher
product
Model choice is becoming a line item, and the differentiator moved up the stack1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 22, 2026