Build1 publisher2 min readPublished
Epoch AI finds GPQA-grade answers getting 13 times cheaper every year
Epoch AI puts the price of answering a GPQA Diamond question at a fixed accuracy bar falling about 13 times a year, faster than compute under Moore's Law. Buyers get that discount only by moving to newer models, so the part of a stack that has to switch cheaply is the evaluation that qualifies each one.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Epoch prices a single GPQA Diamond multiple-choice answer from any model scoring 81.25% or better, so the series follows a capability bar across models.
- OpenAI's o3 cost $0.30 per question in January 2025, and GPT-5.6 Luna matched its 75% score about 18 months later at $0.0004.
- Epoch's AI price series is measured only from 2023, and the points before that are extrapolated from the shorter trend and outside research.
- Chess-puzzle and mathematics benchmarks showed cost efficiency improving at vastly different rates from the GPQA Diamond figure.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- cost A spend pinned for a year at one model's per-question price ends the year paying about 13 times what the same GPQA-grade capability then costs, if the workload tracks the benchmark.
- constraint Teams whose tasks behave like the chess or maths results cannot plan on 13x a year and would have to measure their own cost curve before pricing any commitment.
- exposure Frontier labs selling on a capability lead get a shorter window to profit from it, Tom's Hardware argues, when a cheaper model reaches the same score within months.
The headline numbers agree with each other. Thirteen times a year compounds to about 47% off each quarter [1], matching the report's "just under 50%" [1]. Five years at that rate comes to roughly 371,000x [2]. Epoch's chart shows AI costs falling by hundreds of thousands of times over five years [9].
The worked example falls faster than that. OpenAI's move from o3 to GPT-5.6 Luna is a 750x fall in about 18 months [3]. At the headline rate, 18 months would produce about 47x [4]. The pair also sits below the series that produced the 13x. o3's score falls short of the bar Epoch uses to define that series [3][4].
Epoch compares its line with older technologies [2]. By Tom's Hardware's account, AI's fall runs four times the pace of DNA sequencing and 18 times that of lithium batteries [13]. Lithium batteries fell about 100x between 1991 and 2024 [11]. Some of the comparison series are short. Electricity is tracked only to 1973 [8]. Compute stops in 2001 [8], an odd place to end a comparison with Moore's Law.
Tom's Hardware turns the curve into a provider-loyalty problem. It asks why a customer would stay with a provider once a competitor offers something better and cheaper [10]. The drop it cites happened inside one catalogue: o3 and GPT-5.6 Luna are both OpenAI models [4][5]. The coverage does not break the decline down by vendor.
I think the case for switching holds, but it points at a narrower target. The saving goes to whoever changes models, whether or not the vendor changes. For the 13x to apply at all, a team's workload has to behave like the multiple-choice exam Epoch priced [3]. GPQA Diamond cannot say whether a cheaper model clears a team's own bar on its own tasks.
So the part of the stack that has to be cheap to repeat is qualification. That means a held-out set of real tasks with a pass threshold and a cost per task, re-run each time a cheaper model ships. A provider-neutral client is the smaller job. It pays off only after that evaluation shows another vendor's model passing.
What to watch
- Epoch's published per-discipline rates for chess puzzles and mathematics, and how far each sits from the GPQA Diamond figure.
- Whether the measured post-2023 slope holds once models after GPT-5.6 Luna are priced against the 81.25% bar.
- Whether providers cut prices on existing models to follow their successors; if they do, a buyer pinned to one model shares in the decline without switching.