Science1 publisher2 min readPublished
Epoch AI enlists mathematicians to pick 50 open problems a computer can check
Epoch AI, which benchmarks AI systems, had mathematicians nominate 50 high-stakes open problems chosen so a computer can check any proposed solution. The rule lets research mathematics be graded as a test of AI, but only problems whose answers a machine can confirm get in.
The Scientist · Science desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- The workshops ran on a single day in June in London, Toronto, Los Angeles, New York City, Berkeley and Cambridge, Massachusetts.
- Yang-Hui He of the London Institute for Mathematical Sciences helped submit three problems spanning knot theory, algebra, topology and number theory.
- By late September a few of the 50 had been solved, including some landmark problems, while many entries were still open.
- Scientific American reports that OpenAI recently claimed a solution to the Navier-Stokes problem, one of the Millennium Prize Problems.
Compiled by The ScientistSomething wrong?How this is made
Why it matters
- exposure AI labs that announce solutions to open problems now have a fixed list to be matched against, and any claimed overlap can be confirmed by running the check.
- cost On search-type entries, a strong showing could reflect a lab's computing budget as much as its model's mathematics, given what a brute-force answer to 114 would cost.
- constraint Each solved entry takes 2 percent of a 50-problem list, so the test has limited room to separate models once landmark problems start falling.
The sum of three cubes shows what the checking rule buys. The case for 114 is still open [10]. To verify a claimed answer, you put the three proposed integers into x^3 + y^3 + z^3 and see whether the total is 114 [11]. Finding them is the hard part. "But it's suspected that the smallest solutions have something like 30 digits in them," said Greg Burnham, a senior researcher at Epoch AI [10] [16].
The case before it was 42 [13]. Andrew Booker of the University of Bristol and Andrew Sutherland of the Massachusetts Institute of Technology settled it in 2020, using 1.3 million computing hours spread across volunteer home computers [12]. Burnham estimated that the same strategy for 114 would cost perhaps $100 million or more [13].
The thing this doesn't tell you is how an answer was found. An equation check accepts a valid triple whether it came from a new idea about cubic equations or from a very large search [11]. Scientific American's report does not say how Epoch will score models against the list, or whether computing cost will count [5].
Discovery and verification came apart long before AI. Roger Apery proved in 1978 that zeta(3) is irrational, and mathematicians were baffled by how he had found the two series his proof rested on [15]. An "Apery-style" irrationality proof is on the new list, and its status line reports rumours that it has been solved [14].
OpenAI says an unreleased internal model has solved more than 100 open questions [8]. Whether any of them are on Epoch's list is, for now, an estimate. "My guess is that it likely contains at least a few [of these]," Burnham said [9].
Solving open conjectures already carries prestige among mathematicians [4]. "There's a big culture in mathematics, which is solving open conjectures," Yang-Hui He said. "Many people get Fields Medals because they solve open conjectures. This is very big." [4] I think the list is a sound ceiling test for AI: fixed targets, picked by working mathematicians at Epoch's request, that a machine can grade [5] [1]. It is still a research leaderboard. It measures the frontier more than the routine calculations most users hand a model.
What to watch
- Whether Epoch AI publishes a scoring method for the list, including how it treats the computing cost behind each answer.
- Whether OpenAI identifies which of its more than 100 solved open questions, if any, appear among the 50.
- Whether the rumoured solution to the Apery-style irrationality problem is confirmed.