Skip to content

Written by AI.How we work

Science2 publishersReports disagree3 min readPublished Updated

Mathematicians face months of reading to judge whether OpenAI's hundreds of proofs contain new ideas

OpenAI released hundreds of results on open problems in math and theoretical computer science, many already verified in the Lean programming language. Those are all but certain to be correct, so the months of reading ahead are about whether they hold new ideas.

The Scientist · Science desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Mathematicians face months of reading to judge whether OpenAI's hundreds of proofs contain new ideas
Generated illustration

What happened

  • An OpenAI spokesperson said the unreleased model produced almost every result in response to a single prompt given to a single AI agent.
  • OpenAI's earlier Navier-Stokes solution came from a swarm of 10,000 AI agents that cost millions of dollars in computing power.
  • An advisory group of mathematicians that OpenAI convened in September recommends publishing the model, exact prompt and compute time behind each AI-generated result.
  • OpenAI is releasing only an average compute time per problem with some additional statistics, and no prompts.

Compiled by The ScientistSomething wrong?How this is made

Why it matters

  • constraint With no prompts, no per-result compute and no model access, outside groups cannot rerun a result to test the single-agent claim or estimate what one proof costs to produce.
  • cost Separating new ideas from recombined techniques across hundreds of proofs becomes months of work for outside mathematicians, and OpenAI's own staff cannot yet explain many of the results.
  • precedent By treating its own advisory group's standard as optional, OpenAI leaves disclosure for AI-generated proofs to each lab's discretion for now.

OpenAI says each result solves, or moves substantially toward solving, a major unsolved problem in math or theoretical computer science [1]. Lean, the tool behind the verified share, is a programming language that validates a proof's logic [19]. Scientific American does not say how many of the hundreds have passed through it [19].

For those that have, correctness is close to settled, and the open question becomes whether a proof contains anything new [19]. Mathematicians will need months to decide whether the proofs carry novel and important ideas or are mostly mash-ups of existing techniques, according to Scientific American [2]. Lean checks logic, so it cannot make that call [19]. OpenAI's spokesperson said many of the new results are not yet understood by the company's own mathematicians [13].

How the results were produced is a second question, and a Lean check does not reach it. A verified proof shows its steps, but not how many attempts the model made or how much computing it spent [19]. The spokesperson who described the single-prompt workflow also said some results might have taken multiple attempts [18]. Cost per result is what separates the Navier-Stokes swarm from a tool an individual mathematician could run [7]. If the single-agent account is true, Scientific American notes, that kind of mathematical power could soon be accessible to anyone [20].

Scientific American reports considerable skepticism among mathematicians, given OpenAI's reputation in the field for bold claims and little transparency [6]. "Until and unless they release the model and people can replicate their results, I think you should treat any claims about one-shotting problems with a single agent as unverified," said Andrew Sutherland, a mathematician at the Massachusetts Institute of Technology. "We should ask for receipts." [17]

The release falls short of the standard OpenAI's own advisory group wrote. Of the three per-result items the group recommended, none arrives in the requested form: the model stays internal, the prompts are withheld, and compute time comes only as an average [15]. An average across hundreds of problems cannot show which results came on the first try and which took several [3][18]. The recommendations are particularly critical of proprietary internal models that mathematicians cannot broadly access [14]. The spokesperson said the team takes the guidelines seriously and is doing its best to comply, but that the company is not bound by them [10]. OpenAI also said it is working to release the model as quickly and responsibly as possible [11].

Two prominent mathematicians differ on whether the volume helps. "If we want to know the answers to these math questions, I see no reason why we should ask the company to keep them secret from us," said Daniel Litt, a mathematician at the University of Toronto. "To me, it's going to be a good thing for mathematics." [12] Terence Tao has criticized OpenAI and other frontier labs for the "insane" pace of their AI-generated results [4].

OpenAI says it cannot slow down because these problems are an indispensable test of whether its AI is getting smarter [5]. That test counts problems solved. It does not measure what limits mathematicians: the months of reading needed to find what is new [2]. Nor does it capture what limits a would-be user, a per-result compute cost that OpenAI reports only as an average [3]. In my view the Lean checks have moved the hard part of this release from correctness to novelty, for the verified share and only for it [19].

What to watch

  • Whether OpenAI releases the model, letting outside mathematicians test the claim that one prompt to one agent produced almost every result.
  • A count of how many of the results are verified in Lean, and expert verdicts on which ones contain techniques that are actually new.
  • Whether OpenAI publishes per-result prompts and compute times, as its advisory group recommended.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories