Build1 publisher2 min readPublished
Human mathematicians steered and reviewed Meta's six Muse Spark math papers
Meta's six Muse Spark math papers, out October 2, label which passages the model drafted and which the researchers wrote. They show a chat assistant writing search code and proof drafts for mathematicians who chose the problems.
The Engineer · Build desk
What happened
- The papers name the researchers and their roles, with Aykut Arslan on probability and optimization, Leonard Dinh on differential equations, and Brennan and Golich on group theory.
- Researchers used Muse Spark 1.1 and 1.2 in Thinking Mode through the ordinary Meta AI chat interface, with no custom research tooling, according to Meta.
- In group theory, Muse Spark wrote a GAP search program that found a counterexample, and the researchers checked it before completing the proof themselves.
- Meta announced Muse Spark 1.3 on September 2, and the October papers do not establish that the newer version produced any of these results.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint A team copying this setup needs its own mathematicians to pick problems and verify output, because nothing in the papers shows the model setting an agenda or checking itself.
- contradiction Meta's count of five previously open questions sits beside its admission that other teams reached half the problems independently, so the papers support claims of assistance more than claims of priority.
- precedent Passage-level drafting labels give journals a concrete disclosure standard to ask of later AI-assisted proofs, including Meta's own future papers.
- cost Model-drafted sections still went through researcher checking and revision, so any time saved comes net of expert review hours on every drafted section.
I think the two disproofs are the strongest evidence in the batch of what the model added, because they are the easiest to audit. Brennan and Golich's counterexample is a group with 384 elements [12]. Once a candidate group exists, confirming that it breaks the conjecture is a finite computation, and nobody has to trust the model to run it. The evolution-algebra paper has the same shape. There the model produced a counterexample and proposed alternative characterizations, and a researcher refined them [9].
The proofs ask more of a reader. Dinh's paper proves finite-time blow-up for a specified class of radial solutions to a wave equation, a question left open in 2015 [11]. Arslan's probability paper puts a sharp threshold near n = d^2/4 for whether random Gaussian points can be fit exactly to an ellipsoid [10]. In 100 dimensions that is about 2,500 points [2]. The paper leaves behaviour at the threshold itself unresolved [10]. For the arithmetic-physics paper, which extends a connection first developed for the Tate curve, Meta says the model drafted three technical sections that researchers checked and revised [14].
I'd call the passage labels the best engineering in the release. Each paper marks which passages were drafted primarily by researchers and which by the AI [5]. On a long technical argument, a referee can go straight to the machine-drafted sections and read those hardest.
The novelty claim needs its fine print. Meta says the batch includes proofs addressing five questions it describes as previously open [1]. Its own account acknowledges independent work on three of the six problems [16]. An August 2026 posting covered the ellipsoid threshold [16]. On September 16 the AI agent Nilradical reported a separate counterexample to the group-theory conjecture, so a second AI system also disproved it [16]. Other independent work addressed the evolution-algebra conjecture [16]. Three of six is half the batch [1]. The other teams used different approaches, according to Runtimewire's account of Meta's disclosures [15].
Meta frames the project as a test of whether a general-purpose assistant can help on problems that have no answer key [18]. In these papers the answer key was a second group of mathematicians, separate from the authors, who reviewed the work [4]. The record covers six finished papers, and Meta's account does not say how many problems were attempted along the way [1].
What to watch
- Journal referee decisions on the six papers, especially the three technical sections Meta says the model drafted in the arithmetic-physics paper.
- Any Muse Spark 1.3 math results published with the same passage-level labels, since the current papers cover only versions 1.1 and 1.2.
- Dates from the August ellipsoid posting and Nilradical's September 16 counterexample set against Meta's own timeline for those results.