Build3 publishersIndependently confirmed3 min readPublished
OpenAI hands mathematicians 372 AI-generated proofs to check on GitHub
OpenAI posted 372 results from an internal model to GitHub, at about three hours of ChatGPT Pro compute each on average. Finding proofs is now cheap enough that the slow step is mathematicians checking and judging them.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- OpenAI says nearly every result came from a single prompt to a single agent, though some took more than one attempt.
- Many of the proofs ship with Lean formalizations a machine can check, and OpenAI plans to formalize more.
- The Institute for Advanced Study advisory group OpenAI consulted could advise on how results are communicated, but not on whether or how fast they are produced.
- In an open letter, 25 Fields Medal winners warned that mass-producing true statements could damage mathematics, whose real goal is conceptual understanding.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint Reviewers who trust the Lean files can skip the line-by-line logic check, but only for the formalized subset, and they still have to rule on originality themselves.
- cost OpenAI pays for generation in compute, while mathematicians pay for vetting in their own review time, and the volume could exceed their capacity for manual review.
- precedent If a versioned GitHub release with Lean files is accepted as the record, other labs can circulate machine-generated results without waiting on journal referees.
- contradiction A passing Lean check cannot settle the dispute between OpenAI and the Fields medalists over whether mass-produced results help mathematics or harm it.
Three hours is an average across 372 results [3][1]. OpenAI shared only averages, with no per-problem figures and no prompts [7]. Across the whole set that is roughly 1,116 hours of ChatGPT Pro Thinking compute [22]. Treat it as a number about OpenAI's problem list. For it to transfer to another group's list, their problems would have to resemble the ones this model was given, and failed attempts would have to be counted the same way. Some results took multiple attempts [10]. The Decoder's report does not say whether those retries are inside the average.
The Navier-Stokes run needed a 10,000-agent swarm and millions of dollars [14]. Next to that, one agent per result is a large drop in what a proof costs to find. Checking is slower. The Navier-Stokes solution has been under formal review for weeks, according to OpenAI [15].
Lean is aimed at that gap. It is a programming language built so a machine can check a proof [4]. When the checker accepts a formalized proof, the logic is correct [11]. No referee has to do the line-by-line pass. Lean cannot judge whether a result is relevant or original [11]. Each result is meant to solve or advance an open problem, some tied to the Riemann hypothesis [21]. Whether a given result is new and matters is still a mathematician's call. I think formalization is the right engineering choice for the half of review it covers. I'd want The Decoder's "many" formalized proofs [4] turned into a count before treating the bottleneck as cleared.
The packaging deserves credit. The repository keeps revision logs and citations [2]. OpenAI also published reasoning summaries, statistics on how many problems the model attempted, and compute estimates [6]. A versioned history suits results that will be corrected in public. OpenAI has said it wants to improve its citations and presentation [12]. The Decoder calls the choice of GitHub over journals a statement that the traditional process is too slow for this volume [17].
The doubts reach OpenAI's own advisers. The advisory group it consulted includes Fields Medal winner Timothy Gowers [8][7]. Gowers has warned that within one to two decades the mathematical literature could grow enormously while no human community remains that truly understands it [19]. Terence Tao has argued that training young mathematicians should emphasize the human side and tightly limit AI tool use [20].
OpenAI plans to fund workshops and conferences on understanding AI-produced results [9]. It says it is working on a responsible release of the model to "directly empower scientists with state-of-the-art capabilities" [13]. A wider release would put proof generation at roughly this cost in more hands, with the same human reviewers on the other end.
What to watch
- A count of how many of the 372 results carry Lean formalizations as the planned ones are added.
- The outcome of the formal review of OpenAI's Navier-Stokes solution.
- The terms of OpenAI's planned release of the model to scientists, and whether it comes with per-problem compute figures.