Product1 distinct publisher3 min readUpdated
Timothy Gowers calls OpenAI's results extraordinarily impressive and then sorts them. The sorting, not the praise, tells you which of your problems these models can touch.
The Product Desk · Product desk

Compiled by The Product DeskSomething wrong?How this is made
Timothy Gowers, who won the Fields Medal in 1998, published a blog post on 12 August about the papers behind the current round of AI-solves-mathematics claims, and his central observation is a taxonomy rather than a verdict: the most celebrated results have almost all arrived as counterexamples rather than proofs [1][2][10]. That is not pedantry for anyone being sold these systems, because the two tasks ask for different things from whoever does them.
A proof shows something is always true. A counterexample shows that a claim of always-true fails, by producing one object where it does [12]. Both settle the question. One demands an argument covering every case; the other demands a single object, and permission to keep guessing until you have it [13].
Gowers is not dismissive. He calls the results "extraordinarily impressive" and states that models can prove hard things: "LLMs are not just good at finding counterexamples: they can find proofs of difficult statements as well" [4][5]. OpenAI announced ten solved open problems in mathematics and theoretical computer science, and two led the coverage: the construction of a non-sofic group, which Gowers calls "one of the most important unsolved problems in group theory", and a lower bound showing a multicolour Ramsey number grows superexponentially, which he describes as "a major open problem in Ramsey theory that I didn't necessarily expect to see solved in my lifetime" [6][7][8][9]. His careful summary point is that models prove universal statements perfectly well, but the strongest things they have proved do not match the strongest things they have disproved [11].
He also reclassifies two of the headline items, including one against its own paperwork. On the non-sofic group he notes that several construction routes already existed in the literature and doubts many experts strongly believed all groups were sofic, so he reads it as the first example of a non-sofic group rather than a counterexample, despite OpenAI titling that section of its paper "A counterexample to the soficity conjecture" [14][15]. On the Ramsey result he marks his own homework: many people expected an exponential bound and for them it overturned a belief, but Gowers had worked on an equivalent formulation years ago in what turned out to be the right direction, so for him it confirmed a weak expectation [16].
The operational part is his list of eight ways mathematicians hunt for an example [17]. Four reward volume and broad recall, which is what the machine has: the off-the-shelf check, the step-by-step build, the probabilistic argument and the generic example [18]. Half the list, then, plays to the hardware [23]. Three of the others - leaving parts undefined, trying to prove the opposite, successive approximation - require a judgement about whether the current approach is worth continuing [19]. Gowers locates the gap there and calls it a nose: knowing when you are getting somewhere and when to abandon a branch, which is what lets a human prune a search tree no computer could exhaust [20].
His evidence is partly anecdotal and he says so. Working with GPT-5.6 Pro on open problems, he is often handed approaches that look promising and do not survive scrutiny, including the pattern where the model reports it has reduced the question to a narrower one, "which sounds very promising until it has happened five times without any obvious progress having been made" [21][22].
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Cambridge mathematician Timothy Gowers published a blog post on 12 August on the distinction between counterexamples and proofs in recent AI mathematics results.
Gowers has read the papers, which most people commenting on AI and mathematics have not.
Gowers calls the results "extraordinarily impressive" and says plainly that models can prove hard things too.
Gowers writes: "LLMs are not just good at finding counterexamples: they can find proofs of difficult statements as well."
OpenAI announced ten open problems solved in mathematics and theoretical computer science.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One expert primary read, self-declared partly anecdotal
The strongest evidence is direct: a Fields Medallist who read the papers and attended the talks, quoted at length, giving a structured classification of results and a falsifiable standard. Against that, the cluster is single-publisher, the underlying OpenAI paper is not independently examined here, the capability gap rests on Gowers's own admission that his evidence is partly anecdotal, and no other mathematician is quoted to corroborate the reclassifications.
Frontier results announced and one disclosed research user
Concrete adoption signals are thin and anecdotal: OpenAI's announcement of ten solved problems, and one senior mathematician disclosing hands-on use of GPT-5.6 Pro on open problems. No supplied source gives numbers of researchers using these models, institutional uptake, or any deployment or pricing facts, so the measurement stays low by evidence available rather than by judged insignificance.
Real results, over-labelled
The gap is in labelling rather than substance. Gowers endorses the results as extraordinarily impressive and both headline problems as genuinely major, so the underlying achievement is not overstated. But the framing is: OpenAI's section title claims 'A counterexample to the soficity conjecture' where Gowers reads a first example, few experts strongly held the belief being overturned, and he was neutral on the Ramsey expectation. Coverage also skipped his structural caveat that the strongest proofs do not match the strongest disproofs and that the search-pruning judgement is missing.
Lab framing versus a conceded human-role bias
Incentives pull in both directions and both are visible in the source. OpenAI has an interest in describing its output as counterexamples to conjectures, which is the specific framing Gowers disputes. Gowers in turn concedes he may be clinging to the hope that humans keep contributing for a while, and his falsifiable cap-set standard is set by the same person judging whether it is met. The publisher's angle - the observation 'the aggregators skipped' - also rewards a contrarian read of the announcement.
Credible attribution, single channel
Confidence is moderate: what Gowers said is well documented through direct quotation and he is an unusually qualified source, so the attributed claims are reliable. The wider inferences - that the counterexample skew reflects a durable capability boundary - rest on one expert's partly anecdotal experience, reported by one publisher, with no OpenAI response, no independent expert corroboration, and no verification of the underlying papers in the supplied material.
leadership
Disney swaps raises for discounted stock and a full health-plan re-enrollment1 distinct publisher
science
Text watermarks land on 2 December. The detection they imply does not.1 distinct publisher
invest
A Connecticut judge just priced prompt injection: no fine, no e-filing2 distinct publishers
product
A 2x LLM bill is not a bug report: token spend is an observability problem1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 17, 2026