Skip to content

Product1 publisher3 min readPublished

Hundreds of mathematicians ask policymakers to consult experts on AI math claims

A run of vulnerability, hacking and mathematics announcements from Anthropic and OpenAI drew fast coverage. The expert re-readings that followed described plagiarism accusations and basic security failures.

The Product Desk · Product desk

Illustration accompanying Hundreds of mathematicians ask policymakers to consult experts on AI math claims

What happened

  • An OpenAI-Hugging Face hacking incident followed, after which Anthropic disclosed a similar incident proudly and Meta disclosed one reluctantly, both involving their own models.
  • Two days before OpenAI claimed a mathematical breakthrough of its own, NYU Courant professor Tristan Buckmaster published a statement suggesting OpenAI had stolen other people's work.
  • A statement signed by hundreds of mathematicians asks policymakers to consult experts instead of relying on press releases or popular reporting of mathematical results.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • decision Anyone weighing a model for vulnerability triage has to choose between acting on the announcement and waiting for the specialist reading that, per the column, lands later and with less coverage.
  • exposure Filing an incident as model behavior keeps the vendor's own security practices out of the frame, and the customers running on that vendor inherit the practices.
  • contradiction The column attributes the Astra release to OpenAI while also placing Anthropic's math claim first, leaving unclear which vendor's result the mathematicians actually took apart.

A security lead who read Anthropic's end-of-April claim and wanted to do something with it the following Monday had little to work with. "Better at finding software vulnerabilities than most security experts" is a comparison [1], and the column reporting it does not include a benchmark, a code corpus, or a count of confirmed findings [15]. A team that has to answer for relying on that comparison would need all three to check it itself. In my view the claim as reported justifies a pilot and stops there. What MIT Technology Review describes is a lag between the announcement and the reading. Mathematicians were initially stunned by a press release saying the chatbot Astra had solved problems that "have been open and seen no progress on the main result for at least a decade" [5]. Later they said the results were not as "novel as first appeared" and that Astra had not made a "profound intellectual leap" [6]. Two days before OpenAI claimed a mathematical breakthrough of its own, Tristan Buckmaster, a math professor at New York University's Courant Institute, published a statement suggesting OpenAI had stolen other people's work and improperly attributed it [7]. The second story, the column says, gets less media attention than the first [14]. The column's own account of who published what contradicts itself. It names Astra as OpenAI's chatbot [5] while also saying Anthropic claimed a mathematical breakthrough first and OpenAI followed weeks later [4]. The coverage leaves a buyer guessing which vendor's result the mathematicians examined [18]. There is a practical reason the demos keep landing in mathematics and programming. Answers in those fields can be verified once suggested. That makes the systems easier to tune, because output can be evaluated without paying data workers to annotate each one [11]. The same verifiability sets the ceiling on what the demo tells you. Work whose output cannot be checked by running it, or by one specialist reading it in an afternoon, sits outside the conditions the system was tuned for. The piece describes three capability claims and three incident disclosures from the companies themselves. Those six fell in the roughly 21 weeks between the end of April and the column's publication on 22 September, an average of one every three and a half weeks [16]. At that rate, a team waiting for specialists to publish their reading will have a second announcement to explain before the first one is settled. Two tests sort these announcements before they reach a purchase decision. The first is whether you can check the output yourself, in your own codebase or your own ticket queue, without the vendor's help. The second is whether the vendor published the method, the baseline and the data, so someone outside the company can run it again. A claim that passes both belongs in an evaluation. One that passes only the first is a pilot you fund yourself, and one that passes neither is marketing with a product name in it. Hundreds of mathematicians signed a statement saying there is "currently a strong commercial incentive on the part of the technology industry to overstate the capabilities of their products" [8], and asking policymakers to "consult with experts, including mathematicians, in forming policy decisions rather than relying on press releases or popular reporting of mathematical results" [9]. A procurement committee could take the same instruction. The column argues that the "superintelligence" and "rogue model" labels ascribe agency to products instead of to the companies building them, and help those companies evade accountability [12]. On the hacking incidents, it cites unnamed cybersecurity experts as saying the story is about OpenAI's negligence and its failure to adopt basic, established security practices [3]. Anthropic engineer Jacob Coxon went viral announcing his departure from the company, saying it and OpenAI are "racing straight towards self-improving superintelligence and gambling with our lives" [10].

What to watch

  • Whether Anthropic publishes the corpus and expert baseline behind the Claude Mythos vulnerability claim so a security team can re-run it.
  • Whether OpenAI answers Buckmaster's attribution allegations or the mathematicians' research misconduct and plagiarism accusations.
  • Whether the Sanders bill on artificial superintelligence is redrafted after the mathematicians' request for expert consultation.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories