Skip to content

Invest1 publisher3 min readPublished

18x per joule in 16 months, and most of it was not your model choice

Stanford's Hazy Research group put a decomposed number on local inference efficiency. Hardware did more of the work than architecture, which changes who captures the gain.

The Investor · Invest desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • A paper from Stanford's Hazy Research group titled "Measuring Intelligence Efficiency of Local AI" (arXiv:2511.07885) introduces two metrics for benchmarking local AI inference: Intelligence per Joule (IPJ) and Intelligence per Watt (IPW).
  • The research finds intelligence per joule improved 18-fold in roughly 16 months, from mid-2024 to late 2025.
  • IPJ captures total energy efficiency per task, while IPW reflects how much useful work a system produces for every watt it continuously consumes.
  • Model architecture improvements contributed roughly 3.1x of the IPJ gain.
  • Hardware and accelerator advances delivered a 5.9x boost to IPJ.

Compiled by The InvestorSomething wrong?How this is made

Why it matters

A Stanford Hazy Research paper, "Measuring Intelligence Efficiency of Local AI" (arXiv:2511.07885), introduces two metrics for local inference, Intelligence per Joule and Intelligence per Watt, and reports that IPJ improved roughly 18-fold between mid-2024 and late 2025 [1][2]. That is the first version of this argument with a decomposition attached, which is what makes it usable for a build-versus-buy decision rather than a conference slide.

The split is the part worth reading twice. Model architecture accounted for about 3.1x of the gain and hardware and accelerator advances for about 5.9x [4][5]. Those two multiply to 18.3, so the decomposition is multiplicative and essentially accounts for the whole headline figure with nothing unexplained [6]. In log terms, hardware is about 61 percent of the improvement [7]. Architecture arrives with a download; the paper credits a progression from Mixtral-8x7B to newer mixture-of-experts variants and gpt-oss-120b, which activate only the parameters a query needs [9]. The larger share arrives on a capital expenditure cycle you do not control.

An 18x gain over 16 months implies a doubling roughly every 3.8 months [8]. Any payback model for local inference has to assume the box you buy leaves the efficiency frontier inside two quarters, which argues for shorter depreciation assumptions than most finance teams will volunteer.

Read the hardware list before you extrapolate to your own fleet. The benchmarks span NVIDIA's Quadro RTX 6000, Apple's M4 Max, NVIDIA's B200 and SambaNova's SN40L, and the purpose-built accelerators outperformed the M4 Max on efficiency [10]. So the 5.9x is measured across a range that includes data-center silicon. "Local" in this paper is a spectrum, and the frontier end of it is not sitting under anyone's desk.

The quality claim is similarly bounded: local models reached about 88.7 percent accuracy on single-turn chat and reasoning queries by the end of the evaluation window [11]. Single-turn is the friendly case. Nothing here speaks to long agentic chains or heavy tool use.

The operational finding is the routing one. The paper reports that hybrid local-cloud routing, keeping simple queries on-device and escalating only complex ones, cuts energy, compute and cost by 60 to 80 percent versus running everything in the cloud, without meaningful loss of answer quality [12]. Crypto Briefing's own illustration applies that to a hypothetical 10 million dollar annual cloud inference bill falling to 2 to 4 million [13], which is 6 to 8 million dollars of avoided spend [14]. Treat it as arithmetic on the percentage, not a case study.

One caution on mixing the two metrics. IPW rose 5.3x over a two-year window [15], a doubling roughly every 10 months [16], against IPJ's 3.8 [8]. Different metric, different period; IPJ captures energy per task while IPW captures useful work per watt continuously drawn [3]. Quoting them interchangeably will produce a forecast that is wrong by a factor of two or three.

The profiling harness released with the paper went public around November 2025 [17]. The number that decides this for any given team is not 18x, it is the share of their actual query mix that can stay local at acceptable quality, because that share is what the 60 to 80 percent reduction is really measuring [12]. Run the harness on your own traffic, then look at what your hardware refresh cycle does to a 3.8-month doubling time.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories