Product1 distinct publisher3 min readUpdated
The Verge reports OpenAI published solutions to longstanding open problems, leaving mathematicians shell-shocked. The harder part is what it does to grants and the training of new mathematicians.
The Product Desk · Product desk
Compiled by The Product DeskSomething wrong?How this is made
OpenAI published a set of solutions to longstanding open problems in mathematics, and according to The Verge it landed like a bombshell, setting off a broad argument inside the field [1]. The Verge's London-based AI reporter Robert Hart, who interviewed leading mathematicians for the piece, describes a math world that is shell-shocked and in an existential crisis about what mathematicians are for going forward [2][3].
The timing matters more than the result. Hart dates the shift to the last six months to a year, in which models went from very terrible at this work to, in his words, seemingly genuinely quite good at a professional level [4]. He frames it as mathematics absorbing in that window what other fields have spent roughly five years struggling with [5], which is a compression of somewhere between five and ten times [6].
The capability profile is lopsided in a way that should make anyone cautious about extrapolating. Hart says the models remain truly terrible at arithmetic, at days of the week, and at time, citing Verge reporting by Elissa Welle that ChatGPT could not tell time and still cannot [7]. He checked the old benchmark and found the models can now count the R's in strawberry, though both he and Decoder's host said they suspect that particular case is hard-coded, which neither offered evidence for [8]. Hart's explanation for the gap is that advanced math is largely reasoning rather than calculation, and that academic papers often contain no numbers at all [9]. What the newer systems are good at, he says, is forging connections between areas and applying old methods in new ways [10].
That is the part with budget consequences. The Decoder framing puts it directly: what are academic grants and university programs for, if they exist to train new generations of humans to find and solve outstanding problems and frontier models simply answer those problems [11]. The awkward feature of most research funding is that it buys the search, and the search is exactly the artifact that a published solution set devalues fastest. Verifiable, hard, closed-form problems are the easiest thing for a lab to demonstrate on and the easiest thing for a funder to score. Problem selection, taste, and the unglamorous work of checking machine-produced arguments are none of those things, and they are what is left.
Two cautions before anyone restructures a doctoral program. The first is that a six-to-twelve-month capability move [4] is thin evidence for a five-year pipeline decision, and the podcast is an account of the reaction rather than of independent verification. The second is the possibility Decoder raises itself: that the attention on mathematics is largely a marketing exercise for frontier labs with no stake in the discipline's future [12]. Both readings are consistent with the same publication event.
Watch whether the labs release enough method detail for mathematicians to reproduce the results rather than only react to them, and who absorbs the verification cost if they do not. Watch also whether the claimed transfer of these skills to other domains materializes [13], because that is the assumption most likely to be quietly priced into hiring and grant decisions before it has been tested.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Hart says the change was spurred by a transition in AI capability that exploded in the last six months to a year, in which AI went from being very terrible to seemingly genuinely quite good at a professional level in a very short space of time.
Hart says mathematics is dealing in a very compressed period with a lot of what other fields have been struggling with for the last five years.
OpenAI published a set of solutions to longstanding problems in math that "went off like a bombshell" in the field and caused a huge debate in the math community.
Hart says a lot of advanced math is actually reasoning rather than counting, adding, or multiplying, and that academic math papers often contain no numbers.
Robert Hart is The Verge's London-based AI reporter and spoke to some of the most accomplished mathematicians of our time about OpenAI's published solutions.
Hart says the systems are very good at forging connections between different areas and applying old methods in new ways.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Single secondhand account, primary result uncited
Everything rests on one podcast interview from one publisher. The central factual claim — that OpenAI published solutions to longstanding open problems — is asserted without a link, date, paper identity, or peer-review status, and the mathematicians said to have been interviewed are neither named nor quoted in the supplied excerpt. Capability claims are qualitative ('seemingly genuinely quite good at a professional level') with no benchmark, and the speakers explicitly disclaim two of their own observations (topology weakness unverified; strawberry hard-coding a conspiracy theory). The persistent-failure claims are the best-grounded element, resting on prior Verge reporting plus anecdote.
No adoption evidence supplied
The supplied material reports a publication event and informal spot checks, but gives no evidence of uptake: no count or share of mathematicians using these systems, no institutional deployment, no grant or program decision, no product or pricing action, and no lab disclosure of usage. Debate within a community is not measurable adoption, so no adoption value can be assigned without inferring facts the source does not provide.
Framing runs ahead of verifiable detail
The packaging is maximal — bombshell, shell-shocked, existential crisis, a phase transition implying roughly five-to-ten-times compression of disciplinary adjustment — while the substance inside the same interview is more modest and partly self-undercutting: models still cannot tell time or handle days of the week, the topology weakness cannot be verified, the strawberry explanation is admitted to be a conspiracy theory, the OpenAI result is never cited, and the host concedes the crisis framing is 'pure Decoder bait'. The gap is one of overstated framing rather than false content, so it is positive but moderate, not extreme.
Lab promotion and podcast framing incentives both in play
Incentive pressure is unusually visible because the source names it. The episode explicitly floats that the math attention may be a marketing exercise for frontier labs indifferent to the discipline, which is a live promotional incentive on the OpenAI side that the cluster cannot test. On the publisher side, the host calls the existential-crisis framing 'pure Decoder bait' and the piece cites The Verge's own prior reporting, indicating an engagement and self-referential incentive in how the story is packaged. Score reflects clearly identified incentives on both sides rather than any finding that they distorted the underlying facts.
Low confidence: one publisher, largely unresolved questions
Confidence is constrained by a single-source cluster in which the most consequential material is posed as open questions — grants and training pipelines, cross-domain transfer, and lab motives all end unresolved. The claims that can be relied on are narrow: that The Verge reported the OpenAI publication and framed a disciplinary crisis, and that basic arithmetic and time handling remain weak. Anything about magnitude, mechanism, or institutional consequence would need corroboration outside this cluster.
science
Text watermarks land on 2 December. The detection they imply does not.1 distinct publisher
invest
A Connecticut judge just priced prompt injection: no fine, no e-filing2 distinct publishers
science
Claude's watermark is a compliance artefact, not a cheating detector1 distinct publisher
security
OpenAI's 13-17 tier turns teen AI safety into an age-assurance problem1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 20, 2026