Skip to content

Product1 publisher3 min readPublished

A Fields Medalist read the AI maths papers: most of the wins are counterexamples

Timothy Gowers calls OpenAI's results extraordinarily impressive and then sorts them. The sorting, not the praise, tells you which of your problems these models can touch.

The Product Desk · Product desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Photograph accompanying A Fields Medalist read the AI maths papers: most of the wins are counterexamples
Photo: thenextweb.com

What happened

  • Cambridge mathematician Timothy Gowers published a blog post on 12 August on the distinction between counterexamples and proofs in recent AI mathematics results.
  • Gowers won the Fields Medal in 1998.
  • Gowers has read the papers, which most people commenting on AI and mathematics have not.
  • Gowers calls the results "extraordinarily impressive" and says plainly that models can prove hard things too.
  • Gowers writes: "LLMs are not just good at finding counterexamples: they can find proofs of difficult statements as well."

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

Timothy Gowers, who won the Fields Medal in 1998, published a blog post on 12 August about the papers behind the current round of AI-solves-mathematics claims, and his central observation is a taxonomy rather than a verdict: the most celebrated results have almost all arrived as counterexamples rather than proofs [1][2][10]. That is not pedantry for anyone being sold these systems, because the two tasks ask for different things from whoever does them.

A proof shows something is always true. A counterexample shows that a claim of always-true fails, by producing one object where it does [12]. Both settle the question. One demands an argument covering every case; the other demands a single object, and permission to keep guessing until you have it [13].

Gowers is not dismissive. He calls the results "extraordinarily impressive" and states that models can prove hard things: "LLMs are not just good at finding counterexamples: they can find proofs of difficult statements as well" [4][5]. OpenAI announced ten solved open problems in mathematics and theoretical computer science, and two led the coverage: the construction of a non-sofic group, which Gowers calls "one of the most important unsolved problems in group theory", and a lower bound showing a multicolour Ramsey number grows superexponentially, which he describes as "a major open problem in Ramsey theory that I didn't necessarily expect to see solved in my lifetime" [6][7][8][9]. His careful summary point is that models prove universal statements perfectly well, but the strongest things they have proved do not match the strongest things they have disproved [11].

He also reclassifies two of the headline items, including one against its own paperwork. On the non-sofic group he notes that several construction routes already existed in the literature and doubts many experts strongly believed all groups were sofic, so he reads it as the first example of a non-sofic group rather than a counterexample, despite OpenAI titling that section of its paper "A counterexample to the soficity conjecture" [14][15]. On the Ramsey result he marks his own homework: many people expected an exponential bound and for them it overturned a belief, but Gowers had worked on an equivalent formulation years ago in what turned out to be the right direction, so for him it confirmed a weak expectation [16].

The operational part is his list of eight ways mathematicians hunt for an example [17]. Four reward volume and broad recall, which is what the machine has: the off-the-shelf check, the step-by-step build, the probabilistic argument and the generic example [18]. Half the list, then, plays to the hardware [23]. Three of the others - leaving parts undefined, trying to prove the opposite, successive approximation - require a judgement about whether the current approach is worth continuing [19]. Gowers locates the gap there and calls it a nose: knowing when you are getting somewhere and when to abandon a branch, which is what lets a human prune a search tree no computer could exhaust [20].

His evidence is partly anecdotal and he says so. Working with GPT-5.6 Pro on open problems, he is often handed approaches that look promising and do not survive scrutiny, including the pattern where the model reports it has reduced the question to a narrower one, "which sounds very promising until it has happened five times without any obvious progress having been made" [21][22].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories