Leadership1 distinct publisher3 min readPublished
Stanford's AI Index puts a 142-fold parameter cut and a more than 280-fold price cut behind one fixed MMLU threshold. That narrows where building your own still pays, and it lands in a state-law count that doubled in a year.
The Board Room · Leadership desk

Compiled by The Board RoomSomething wrong?How this is made
Two of the Index's curves run in opposite directions, and the distance between them can be worked out. Training compute for notable models doubles roughly every five months [5], which is 2.4 doublings a year, or about 5.3 times more compute. Machine learning hardware costs fall about 30% a year [6], so each unit of that compute costs about 0.7 of last year's price. Multiply the two and a run that stays at the frontier costs about 3.7 times more each year [16]. Cheap tokens and expensive training runs are one market seen from opposite ends.
Which end you stand on depends on whether your quality bar moves. Both headline reductions are pinned to a fixed level: 60% on MMLU for the parameter count [1], and GPT-3.5-equivalent accuracy of 64.8% for the price [3]. What has become nearly free is a 2022 grade of capability. For work whose bar was set then and has not moved, the build case is better than it was; for work that needs whatever is best this quarter, the arithmetic above governs.
That price fell more than 280-fold in eighteen months [3], which could argue for waiting rather than building, since a further drop might land before any in-house project ships. But inference prices have fallen between 9 and 900 times a year depending on the task [4]. A rate uncertain across two orders of magnitude offers no planning input; the level observed at a point in time is what can actually be used. The most recent price point in this record is October 2024 [3], in a report published on 7 April 2025 [12].
The board-deck version says unit costs collapsed, organisational AI use went from 55% to 78% of survey respondents in a year [19], and the response is to insource. It is incomplete in two places. Two-thirds of the 149 foundation models released in 2023 carried open weights [13], which is the part of the build case that holds, but on RE-Bench top systems beat human experts fourfold on two-hour tasks and lose two to one at 32 hours [11]. What is cheap is short work.
The legal side moved without waiting for either decision. Stanford's tally reports how many state AI laws passed, not what they require, and federal bills rose while passage stayed low [9]. Reported incidents reached 233 in 2024 [10], up from about 149 the year before [17]. Those are counts, so anyone sizing a compliance line is estimating from outside the Index, and the obligation follows the deployment: training your own model does not change which state a user sits in.
The decision this quarter is therefore narrower than the price curve implies. The buy case holds for anything measured against the current frontier, and the build case holds where the bar is fixed and the tasks are short. The consequence to name for next quarter is that both routes answer to the same state-by-state surface [9], and this record gives no figure for what answering it costs.
Ranked by verification strength, evidence, and original report placement.
The cost of querying an AI model scoring the equivalent of GPT-3.5 (64.8% accuracy) on MMLU dropped from $20.00 per million tokens in November 2022 to $0.07 per million tokens by October 2024 (Gemini-1.5-Flash-8B), a more than 280-fold reduction in about 18 months.
Depending on the task, LLM inference prices have fallen anywhere from 9 to 900 times per year.
In 2022, the smallest model registering a score higher than 60% on the MMLU benchmark was PaLM, with 540 billion parameters.
By 2024, Microsoft's Phi-3-mini, with 3.8 billion parameters, reached the same 60% MMLU threshold, a 142-fold reduction in over two years.
Training compute for notable AI models is doubling approximately every five months.
Machine learning hardware price performance has improved, with costs dropping 30% per year; 16-bit floating-point performance has grown 43% annually and energy efficiency 40% annually.
Distinct publishers with included, body-backed reporting in this cluster.
2 articles · September 5, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
product
Washington's secret AI test is coming for open weights, and release dates go with it2 distinct publishers
product
Model choice is becoming a line item, and the differentiator moved up the stack1 distinct publisher
invest
Alphabet stock rises 0.6% the same day judge rejects forced sale of its ad exchange3 distinct publishers
leadership
Nvidia is buying the model hub that Intel, AMD and Amazon helped fund4 distinct publishers
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One institution, quoted twice
Every number here traces to a single annual report, reached through two of its own pages, and no outside count appears anywhere in our coverage. The report is precise where precision matters: models named, price points dated, the incidents database identified, a stated publication date of 7 April 2025. What holds the score down is that the corroboration is internal, the arithmetic checks that pass (540 divided by 3.8 gives 142, $20.00 divided by $0.07 gives 286) are ours rather than a second party's, and the report's own interval label of roughly 18 months describes a 23-month span.
Priced, shipped, mostly self-reported
The cheap end of the curve attaches to a product with a date, Gemini-1.5-Flash-8B at $0.07 per million tokens in October 2024, and the small-model claim to Phi-3-mini at 3.8 billion parameters, so the capability described is buyable rather than announced. Past those two anchors the evidence turns into aggregates: 78% organisational use from a survey, 131 state laws, 233 catalogued incidents, 149 foundation models in 2023 with two-thirds released as open weights. Those are real signals, though none is traceable to a named deployment.
Cheap at a 2022 bar
The headline claim is accurate and narrower than it sounds. Both the 142-fold parameter cut and the 286-fold price cut are measured against a 60% MMLU threshold set in 2022, so they price yesterday's capability, while the same report has frontier training compute doubling every five months. The stated pace is also flattered by the interval: November 2022 to October 2024 runs about 23 months, not the 18 both pages give, which trims roughly a fifth off the implied annual rate.
University imprint, mixed committee
The publisher is a university institute rather than a vendor, and says so on the page: an independent initiative at Stanford HAI, steered by a committee of academics and industry people. That composition is worth holding next to the report's own finding that nearly 90% of notable 2024 models came from industry, since the organisations producing the models sit close to the group deciding which indicators get charted. Neither page names a funder, sponsor or data contributor.
Firm figures, single voice
The individual numbers are stated crisply enough to argue with, and the two arithmetic identities we can test hold. The limits are structural rather than sloppy: one publisher, one report, no replication, a survey under the adoption figure, and an elapsed-time inconsistency that neither page catches. That is enough to quote the index confidently and not enough to treat any single figure as settled measurement.