Skip to content

Invest2 publishersIndependently confirmed3 min readPublished

Lean proof checks cover about 300 of the 722 math manuscripts OpenAI's unreleased model produced

OpenAI published 722 math manuscripts across 372 result families, all produced by an internal model it has not released. About 300 come with Lean formalizations a machine can check, Crypto Briefing reported, so the rest wait on human reviewers whom one mathematicians' group is urging to stop working with OpenAI.

The Investor · Invest desk

How we use AISend a correction

What happened

  • OpenAI has not shared the complete prompts used to generate the results.
  • Two days after the release, the Association for Human Mathematics urged mathematicians to stop working with OpenAI.
  • AHM says the Institute for Advanced Study advisory group OpenAI consulted opened by saying AI labs should not test advanced math problems on internal models.

Compiled by The InvestorSomething wrong?How this is made

Why it matters

  • constraint Without the model or full prompts, outside reviewers can grade the selected outputs but cannot measure the model's record across everything it attempted.
  • cost Roughly 420 manuscripts with no machine check need human reading time, and mathematicians pay that cost while OpenAI keeps the system that produced them.
  • contradiction OpenAI cites the Institute for Advanced Study group as guidance for the project, while AHM says that group's opening position opposed running this kind of work on internal models.
  • decision Individual mathematicians now have to choose between triaging OpenAI's catalog and siding with a 752-member association that asks them to cut ties with the company.

OpenAI decided what the field gets to grade. According to Crypto Briefing, the model was run against about 4,000 problems, and the catalog is the subset of outputs the company considered significant [10]. Set the 372 result families [1], or the more than 300 problems The Washington Post counted [19], against that 4,000 and the catalog covers between 7.5% and 9.3% of what was attempted [21]. Neither account reports how the model did on the rest.

Inside the catalog, checking splits in two. Lean is a proof assistant that checks every logical step of a proof the way a compiler checks code [9]. About 42% of the results, around 300, had Lean formalizations as of early October [8], and for those a machine has already confirmed that each step follows from the one before [9]. The other 420 or so manuscripts need a mathematician to read them [22]. Three retractions for errors shortly after release [11] come to about 0.4% of the catalog [23], counted in the first days of reading.

The Lean-checked results could hold and get cited and built on by human researchers, the test Crypto Briefing says mathematics ultimately applies [17]. Retractions could instead climb as readers reach the unformalized 420 [22]. Or part of the field could decline to read at all, and much of the catalog would sit unchecked.

I think review can settle the outputs and cannot settle the claim about the model. OpenAI has shared neither the underlying model nor the complete prompts [3], so a reader can confirm a proof but cannot rerun what produced it. The checking cost moves to mathematicians while OpenAI keeps the system. The counter-case is strong. A correct proof is correct whoever wrote it, and OpenAI says the work solved or made significant progress on hundreds of open problems [7]. The release also stops short of claiming any Millennium Prize Problem fully resolved [18]. If the unformalized results survive reading with few further retractions and OpenAI publishes more on its methods, the model question matters less than I am arguing.

September's Navier-Stokes announcement, reportedly about 10,000 AI agents working over 88 hours [2], or 880,000 agent-hours [20], drew criticism that OpenAI put speed ahead of independent verification [16]. This time the company added a GitHub repository with protocols for revisions and citations [12] and guidelines from an Institute for Advanced Study advisory group that drew on feedback from hundreds of mathematicians [4].

The Association for Human Mathematics says that group opened its first advisory statement by saying frontier AI companies should not test advanced mathematical problems on their internal models [5]. "We urge mathematicians to discontinue their work with OpenAI and to return to a vision of science that centres human understanding," AHM said [14]. Its 752 reported members could each take one manuscript and leave 30 people over [24]. They have pledged instead not to work with commercial AI companies or publish AI-generated mathematical writing [15].

What to watch

  • How many further retractions come out of the roughly 420 manuscripts that lack Lean formalizations.
  • Independent verification of the Riemann hypothesis-related result, which the release does not claim as a solution.
  • Whether OpenAI discloses more about its methods, prompts or the outcomes on problems outside the catalog.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence50
Adoption
Insufficient
Hype gap+35
Incentives70
Confidence50
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    On October 6-7, 2026, OpenAI published 722 manuscripts organized into 372 result families, covering number theory, algebraic geometry, analysis and theoretical computer science, generated by an internal model it has not released.

  2. [2]

    In September 2026 OpenAI reported progress on the Navier-Stokes equations, an effort that reportedly involved about 10,000 AI agents working over 88 hours.

    ReportedSupportedSource: Crypto Briefing, citing reports2 sources— create a free account to open themView cited source
  3. [3]

    OpenAI has not shared the underlying model or the complete prompts used to generate the results, prompting concerns about verifiability.

Sources

2 independent publishers whose own reporting we read for this story.

  1. cryptobriefing.com

    1 article · October 8, 2026

    Mathematicians are combing through OpenAI’s AI-generated results on over 300 problems
  2. indianexpress.com

    1 article · October 8, 2026

    ‘Mathematicians did not ask for this’: Math group calls on researchers to reject OpenAI

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Loading related stories