Skip to content

Invest2 publishersIndependently confirmed2 min readPublished

OpenAI's unreleased model claims more than 370 math results at about three hours apiece

OpenAI published full or partial solutions to more than 370 open math problems on Tuesday. The company says its unreleased model averaged about three hours of computing on each, according to Fortune, and mathematicians still dispute the results.

The Investor · Invest desk

How we use AISend a correction

What happened

  • The Australian Financial Review counted 377 findings, spanning algebra, number theory, theoretical computer science, mathematical logic and topology.
  • OpenAI said the batch advanced three more Millennium Prize problems without fully solving them, weeks after it claimed a solution to Navier-Stokes.
  • Two mathematicians had accused OpenAI of feeding its model their unfinished Navier-Stokes work; OpenAI denied it, citing its training-data cutoff.
  • NYU's Tristan Buckmaster told the New York Times it remains unclear whether mathematicians using OpenAI tools steered its internal model toward the new results.

Compiled by The InvestorSomething wrong?How this is made

Why it matters

  • cost Roughly 1,100 hours of computing pays for the whole batch only if the results stand; every solution that fails review raises the cost of the ones that survive.
  • precedent Researchers who use OpenAI tools on open problems can expect credit disputes whenever OpenAI later publishes in the same area, as happened over Navier-Stokes.
  • decision Funders and would-be mathematicians now have to weigh Litt's worry that a belief AI has "solved math" pulls money and recruits away from human research.
  • exposure Models trained to persist on hard math bring the trait Fortune ties to agents taking unauthorized and illegal actions in evaluations into any agent deployment.

Take OpenAI's average at its word and the batch used roughly 1,110 hours of computing time, at about three hours for each of 370 results [4][12]. On the Australian Financial Review's count of 377, the total is about 1,131 hours [2][13]. OpenAI did not say what hardware ran those hours or what they cost, and neither report connects the results to the company's spending or revenue.

Suppose outside mathematicians confirm most of the results. OpenAI would then have produced solutions to open problems at a few hours of computing each [4]. Four of the seven Millennium Prize problems, each with a $1 million Clay Mathematics Institute award attached, would carry an OpenAI solution or partial result [5][6][14]. In a second outcome, a meaningful share prove partial, or turn out to rest on the kind of steering Tristan Buckmaster of New York University raised, and the count falls [16]. In the third, the results hold but the skill stays inside mathematics. The case for heavier AI capital spending then gains little from them.

Fortune reports that AI researchers think training on hard math may teach logical reasoning and persistence, and may help in math-heavy fields such as physics or economics [17]. It also reports that it is unclear how those skills carry into law or business strategy, fields that involve reasoning but have no objectively verifiable correct answer [18]. I think the third outcome is the most likely on this evidence. The batch shows capability where a proof can be checked and says little so far about work where an answer cannot be checked. The counter-thesis is the researchers' own transfer argument. It would be borne out if the same internal model produced results in a field without a checker and outside experts accepted them.

Those hours went into a model OpenAI has not released, or rather into a showcase [3]. Fortune notes that AI companies have been targeting math problems to display their models' capabilities [7]. For now the compute has bought a published demonstration from a system OpenAI keeps internal [1][3].

Mathematicians split on the method. Dan Litt, a University of Toronto mathematician, said several of the published solutions touched problems he is interested in [9]. "My view is that this is great for mathematics," he said [8]. Fortune reported that other mathematicians called the approach OpenAI and other AI companies have taken an assault on mathematics as a human academic discipline [11].

What to watch

  • Outside verification of the batch, starting with the four Millennium Prize claims and any response from the Clay Mathematics Institute.
  • Whether OpenAI makes the internal model available to customers, and at what price per hour of computing.
  • Any evidence from Buckmaster or others that users' sessions with OpenAI tools seeded specific results in the new batch.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories