Build2 publishersAlso reported elsewhere2 min readPublished
One sign error pulled three of OpenAI's AI-generated math papers within a day of release
OpenAI withdrew three AI-generated math papers on October 7, one day after releasing 722 manuscripts, because of a sign error in an argument they shared. Only 300 of the 719 remaining results have Lean proofs a computer can check, so most of the collection still depends on human review.
The Engineer · Build desk
What happened
- The error sat in a stabilization-trace cancellation argument, and OpenAI says it also broke the construction that two dependent papers relied on.
- OpenAI says it gave its unreleased internal model about 4,000 problems, then grouped the results into manuscripts and families after judging their significance.
- On September 29, an independent advisory group OpenAI consulted urged labs to stop testing advanced math problems on proprietary models and to fund work on human understanding.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- cost The remaining expense falls on whoever reviews the 419 results without Lean proofs, and that effort sits outside OpenAI's three-hour-per-result compute figure.
- exposure Results built on shared constructions fail together, so one unchecked error in a reused argument can take its dependent papers down with it.
- constraint Outside mathematicians can audit the outputs but cannot rerun the unreleased model, so verification has to work from the manuscripts and 10 reasoning summaries alone.
- precedent A public, versioned repository with withdrawal notices and archived copies gives other labs a working template for retracting AI-generated results quickly.
Add the three withdrawals to the 14 revised papers and the 13 with reference updates, and the change history touches 30 manuscripts [19]. That is about 4 percent of the 722 OpenAI first published [20]. The reporting ties the sign error to the three withdrawals [5]. The other 27 entries are repaired proofs, corrected statements, clarified assumptions and reference updates [7].
In my view the repository handled the correction well. It is public and versioned, so the withdrawal shows up in the change history, and each withdrawn manuscript carries a notice linking to its archived version [6]. A reader who cited one of the three can still see what it said [6].
Lean is the machine check. It is a programming language for encoding proofs so a computer can verify them [10]. About 42 percent of the top-line results have a Lean formalization [9]. That leaves 419 of the 719 without one [18]. The README is candid about this: results sit at different verification stages, and unformalized work could contain issues [8].
Generation has a published cost. OpenAI says the average retained result used compute equal to roughly three hours of ChatGPT Pro thinking [11]. Across 719 results that comes to about 2,150 hours [21], in a unit of measure with one supplier. Runtimewire notes that the figure is OpenAI's own compute comparison, not a measure of human labor saved or mathematical value produced [11].
The reporting does not say how the sign error was found or how many reviewer hours the corrections took. On this evidence, verification is the unfinished part of the job [8]. The published figures measure generation only, so they cannot show that verification is also the larger cost [11].
The selection step adds a second layer to check. Runtimewire points out that OpenAI's process of assessing significance and grouping results describes how the catalog was assembled, and does not independently establish that each remaining result is correct, novel or useful [17].
The advisory group OpenAI consulted has no decision-making power at any AI company [13]. It warns that a person who prompts an AI can end up holding a mathematical argument they cannot follow, check or answer for, and it says the mathematical community must assess the work [16]. OpenAI's October 6 post says the company plans to fund workshops, conferences and special programs on understanding AI-generated results [15].
What to watch
- Further withdrawals or revisions in the change history, especially among the 419 results that have no Lean formalization.
- Whether the count of Lean-formalized results rises from 300, and how fast.
- Whether OpenAI names or releases the model, which would let outside researchers reproduce results from the same system.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+35
- Incentives62
- Confidence55
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
OpenAI withdrew three AI-generated mathematical manuscripts on October 7 after finding a sign error that invalidated an argument used across the papers, according to the repository's change history.
- [2]
The withdrawal came one day after OpenAI published a collection it initially described as 722 manuscripts in 372 families of results, produced by an unreleased internal model.
- [4]
OpenAI's October 6 research post says the work was generated by an internal frontier model; it does not name the model or make it available to other researchers.
- [5]
The sign error was in a stabilization-trace cancellation argument; OpenAI says the error also affected the construction used by two dependent papers.
ReportedSupportedSource: OpenAI, via repository change history as reported by runtimewire.comView cited source - [6]
The withdrawn manuscripts carry notices linking to archived versions, and the repository's public, versioned record makes the correction visible.
- [7]
The change history records revisions to 14 other papers, including repaired proofs, corrected statements and clarified assumptions, as well as updates to references in 13 additional manuscripts.
- [8]
The repository README says the results are at different verification stages, that some lack Lean formalizations and that unformalized work could contain issues.
- [9]
OpenAI's repository says 300 of the 719 top-line results have Lean formalizations, about 42 percent.
- [10]
Lean is a programming language used to encode proofs so a computer can check them.
- [11]
OpenAI says the average retained result used compute equivalent to roughly three hours of ChatGPT Pro thinking; runtimewire notes this is OpenAI's compute comparison, not a measure of human labor saved or mathematical value produced.
- [12]
The release includes 10 summaries of model reasoning; the underlying model remains unreleased, limiting outside researchers' ability to reproduce the work from the same system.
- [13]
OpenAI said it consulted the independent Advisory Group on Mathematics and Artificial Intelligence; the group says it operates independently of AI companies and has no decision-making power at any of them.
- [14]
The advisory group's September 29 guidance urged AI labs to stop testing advanced mathematical problems on proprietary models, called for timely release of results and supporting materials, and said labs should fund work needed to build human understanding.
- [15]
OpenAI's October 6 post says it plans to fund workshops, conferences and special programs around understanding AI-generated results.
- [16]
The advisory group says AI can produce mathematical arguments without the person who prompted it being able to understand, verify or take responsibility for them, and that the mathematical community must assess the work.
- [17]
OpenAI says the model was given about 4,000 problems, with results grouped into manuscripts and families after OpenAI assessed their significance; runtimewire says that process does not independently establish that each remaining result is correct, novel or useful.
- [18]
419 of the 719 top-line results have no Lean formalization.
- [19]
The change history touches 30 manuscripts: three withdrawn, 14 revised, 13 with reference updates.
- [20]
The 30 touched manuscripts are about 4 percent of the 722 originally published.
- [21]
Across the retained results, OpenAI's compute comparison totals about 2,150 hours of ChatGPT Pro thinking equivalent, assuming the 719 current top-line results are the retained set.
Sources
2 independent publishers whose own reporting we read for this story.
- cryptobriefing.comOpenAI pulls three AI-generated math papers one day after release
1 article · October 8, 2026
- runtimewire.comOpenAI withdraws three math papers after a sign error breaks linked arguments
1 article · October 8, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.
Topics
- Formal VerificationFollow
- Research ReproducibilityFollow
- AI-Generated MathematicsFollow
Entities
- OpenAIFollow
- Sam AltmanFollow
- LeanFollow
- Advisory Group on Mathematics and Artificial IntelligenceFollow
- ChatGPT ProFollow
- LooptFollow