Invest2 publishersIndependently confirmed3 min readPublished
Lean proof checks cover about 300 of the 722 math manuscripts OpenAI's unreleased model produced
OpenAI published 722 math manuscripts across 372 result families, all produced by an internal model it has not released. About 300 come with Lean formalizations a machine can check, Crypto Briefing reported, so the rest wait on human reviewers whom one mathematicians' group is urging to stop working with OpenAI.
The Investor · Invest desk
What happened
- OpenAI has not shared the complete prompts used to generate the results.
- Two days after the release, the Association for Human Mathematics urged mathematicians to stop working with OpenAI.
- AHM says the Institute for Advanced Study advisory group OpenAI consulted opened by saying AI labs should not test advanced math problems on internal models.
Compiled by The InvestorSomething wrong?How this is made
Why it matters
- constraint Without the model or full prompts, outside reviewers can grade the selected outputs but cannot measure the model's record across everything it attempted.
- cost Roughly 420 manuscripts with no machine check need human reading time, and mathematicians pay that cost while OpenAI keeps the system that produced them.
- contradiction OpenAI cites the Institute for Advanced Study group as guidance for the project, while AHM says that group's opening position opposed running this kind of work on internal models.
- decision Individual mathematicians now have to choose between triaging OpenAI's catalog and siding with a 752-member association that asks them to cut ties with the company.
OpenAI decided what the field gets to grade. According to Crypto Briefing, the model was run against about 4,000 problems, and the catalog is the subset of outputs the company considered significant [10]. Set the 372 result families [1], or the more than 300 problems The Washington Post counted [19], against that 4,000 and the catalog covers between 7.5% and 9.3% of what was attempted [21]. Neither account reports how the model did on the rest.
Inside the catalog, checking splits in two. Lean is a proof assistant that checks every logical step of a proof the way a compiler checks code [9]. About 42% of the results, around 300, had Lean formalizations as of early October [8], and for those a machine has already confirmed that each step follows from the one before [9]. The other 420 or so manuscripts need a mathematician to read them [22]. Three retractions for errors shortly after release [11] come to about 0.4% of the catalog [23], counted in the first days of reading.
The Lean-checked results could hold and get cited and built on by human researchers, the test Crypto Briefing says mathematics ultimately applies [17]. Retractions could instead climb as readers reach the unformalized 420 [22]. Or part of the field could decline to read at all, and much of the catalog would sit unchecked.
I think review can settle the outputs and cannot settle the claim about the model. OpenAI has shared neither the underlying model nor the complete prompts [3], so a reader can confirm a proof but cannot rerun what produced it. The checking cost moves to mathematicians while OpenAI keeps the system. The counter-case is strong. A correct proof is correct whoever wrote it, and OpenAI says the work solved or made significant progress on hundreds of open problems [7]. The release also stops short of claiming any Millennium Prize Problem fully resolved [18]. If the unformalized results survive reading with few further retractions and OpenAI publishes more on its methods, the model question matters less than I am arguing.
September's Navier-Stokes announcement, reportedly about 10,000 AI agents working over 88 hours [2], or 880,000 agent-hours [20], drew criticism that OpenAI put speed ahead of independent verification [16]. This time the company added a GitHub repository with protocols for revisions and citations [12] and guidelines from an Institute for Advanced Study advisory group that drew on feedback from hundreds of mathematicians [4].
The Association for Human Mathematics says that group opened its first advisory statement by saying frontier AI companies should not test advanced mathematical problems on their internal models [5]. "We urge mathematicians to discontinue their work with OpenAI and to return to a vision of science that centres human understanding," AHM said [14]. Its 752 reported members could each take one manuscript and leave 30 people over [24]. They have pledged instead not to work with commercial AI companies or publish AI-generated mathematical writing [15].
What to watch
- How many further retractions come out of the roughly 420 manuscripts that lack Lean formalizations.
- Independent verification of the Riemann hypothesis-related result, which the release does not claim as a solution.
- Whether OpenAI discloses more about its methods, prompts or the outcomes on problems outside the catalog.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence50
- Adoption
- Insufficient
- Hype gap+35
- Incentives70
- Confidence50
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
On October 6-7, 2026, OpenAI published 722 manuscripts organized into 372 result families, covering number theory, algebraic geometry, analysis and theoretical computer science, generated by an internal model it has not released.
ReportedSupportedSource: Crypto Briefing2 sources— create a free account to open themView cited source - [2]
In September 2026 OpenAI reported progress on the Navier-Stokes equations, an effort that reportedly involved about 10,000 AI agents working over 88 hours.
ReportedSupportedSource: Crypto Briefing, citing reports2 sources— create a free account to open themView cited source - [3]
OpenAI has not shared the underlying model or the complete prompts used to generate the results, prompting concerns about verifiability.
ReportedSupportedSource: Crypto Briefing2 sources— create a free account to open themView cited source - [4]
An advisory group from the Institute for Advanced Study provided guidelines for how the project was run, incorporating feedback from hundreds of mathematicians.
ReportedSupportedSource: Crypto Briefing2 sources— create a free account to open themView cited source - [5]
AHM said the Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study, which OpenAI consulted, opened its initial advisory statement by saying frontier AI corporations should not test advanced mathematical problems on their internal models.
ReportedSupportedSource: AHM statement, reported by Indian Express2 sources— create a free account to open themView cited source - [6]
According to AGMAI's recommendations, the use of proprietary internal models by AI labs for mathematical research risks creating a two-tier system where labs outrun the rest of the field.
ReportedSupportedSource: Indian Express2 sources— create a free account to open themView cited source - [7]
OpenAI said the work included solutions to, or significant progress on, hundreds of open problems in mathematics and theoretical computer science.
ReportedSupportedSource: OpenAI, reported by Indian Express2 sources— create a free account to open themView cited source - [8]
As of early October 2026, 42% of the results, around 300, had Lean formalizations.
- [9]
Lean is a proof assistant that checks every logical step of a proof the way a compiler checks code; if a proof compiles in Lean, a machine has confirmed that each step follows from the previous one.
- [10]
The model was tested against approximately 4,000 problems; the released catalog is the subset of outputs OpenAI considered significant.
- [11]
Three manuscripts were retracted shortly after release due to errors.
- [12]
OpenAI published the results in a GitHub repository with protocols for paper revisions and citations.
- [13]
Two days after OpenAI released the results, the Association for Human Mathematics issued a statement criticising the release and urging mathematicians to discontinue their work with OpenAI.
- [14]
"We urge mathematicians to discontinue their work with OpenAI and to return to a vision of science that centres human understanding."
- [15]
AHM reportedly has 752 members, who pledge not to work with commercial AI companies or publish AI-generated mathematical writing.
- [16]
After its Navier-Stokes claim, OpenAI was criticised for prioritising speed and AI-generated results over independent mathematical verification and transparency.
- [17]
Whether results get cited and built upon by human researchers is ultimately how mathematics decides what counts.
- [18]
OpenAI did not claim complete resolutions for any Millennium Prize Problems in the release; a result connected to the Riemann hypothesis has drawn particular attention, and nothing in the release claims it is solved.
- [19]
According to The Washington Post, mathematicians are analyzing OpenAI's findings on over 300 problems.
- [20]
The Navier-Stokes effort amounted to about 880,000 agent-hours.
- [21]
The released catalog covers between about 7.5% and 9.3% of the roughly 4,000 problems tested.
- [22]
About 420 of the 722 manuscripts lack Lean formalizations.
- [23]
Three retractions equal about 0.4% of the 722 manuscripts.
- [24]
AHM's 752 reported members exceed the 722 manuscripts by 30.
Sources
2 independent publishers whose own reporting we read for this story.
- cryptobriefing.comMathematicians are combing through OpenAI’s AI-generated results on over 300 problems
1 article · October 8, 2026
- indianexpress.com‘Mathematicians did not ask for this’: Math group calls on researchers to reject OpenAI
1 article · October 8, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.
Topics
- Formal VerificationFollow
- AI for MathematicsFollow
- Research transparencyFollow