Product1 publisher3 min readPublished
OpenAI hands its new math advisory panel a backlog of results from an unreleased model
OpenAI's new mathematicians' panel must first help release scores more results the company says an unreleased model produced. Researchers call the panel a good first step, though many are unsure how much say it will have over how OpenAI announces what its models find.
The Product Desk · Product desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- Mathematicians who spoke to The Verge, one panel member among them, described the panel's rollout as messy and confusing.
- OpenAI's claimed solution to a Millennium Prize problem drew unease from many mathematicians, who see the company disregarding the field's long-standing norms.
- Thom has since accused OpenAI of "dishonest" behavior, saying ChatGPT chats with him and colleagues may have fed the company's success.
Compiled by The Product DeskSomething wrong?How this is made
Why it matters
- decision OpenAI now has to set a release pace for its unreleased model's results, trading a run of first-claim headlines against giving specialists time to check each one.
- exposure Mathematicians who work through open problems in ChatGPT cannot currently tell whether those chats end up inside a lab's announced result, the gap Thom is pressing on.
- precedent Because the panel's brief extends to other AI companies, the terms it sets for OpenAI's backlog are likely to become the yardstick for Anthropic and any other lab claiming math results.
- cost A bulk release puts the review burden on the specialists who would have to check scores of new results in their own subfields at the same time.
Andreas Thom works on non-sofic groups. One of the 10 results OpenAI announced last month sits in that area, and OpenAI acknowledged it built heavily on earlier work by Thom and Gábor Kun [13]. In a series of Mastodon posts, Thom raised the concern that conversations he and colleagues had with ChatGPT before the announcement may have helped the model get there [11]. According to The Verge, he accused OpenAI of unethical and "dishonest" behavior and of a lack of transparency about where its training data comes from [12]. His posts came days after a separate row over whether the company's models had benefited from unpublished work [10].
Thom's posts are about the conversations [11]. On his account, researchers take problems in their own field to ChatGPT before any result is public [11]. Judging by OpenAI's acknowledgment, the lab assumes that crediting the prior published papers settles the model's debts [13].
On the work itself, the labs have delivered. OpenAI, Anthropic and others announced results on long-standing problems this past year, some beyond what researchers expected current systems could do, including a resolution of one of the Millennium Prize problems [1]. The Verge wrote that OpenAI has shown it can "make impressive breakthroughs in mathematics, then colossally screw up announcing them" [3]. Results that might normally have been celebrated sparked backlash instead [2].
The new panel is aimed at the second half of that quote. It advises OpenAI and other AI companies on how they deal with mathematical research, including how results are presented and released [4]. Its first assignment is the queue. The panel is to help coordinate the release of scores more results that OpenAI says an unreleased model has produced, and the prospect is already stirring dread among researchers, according to The Verge [7].
Whether the panel matters depends on what powers it has. Researchers told The Verge it is a good first step. Many also asked what it will actually do, how much influence it will have, whether OpenAI will listen, and whether a small group of prominent researchers can speak for the wider field [8]. Its brief, as The Verge reports it, is advisory, and the reporting does not describe any power to hold back a release [4].
In my view the panel will be judged on one call OpenAI still controls: the pace of the backlog. Releasing those results at the speed specialists can check them, with provenance attached to each, costs OpenAI a run of first-claim headlines. Releasing them in bulk repeats the pattern that produced the backlash, and the mathematicians already dreading the volume absorb it [2][7]. AI labs say they are learning from their earlier mistakes [14].
The test I'd apply to any AI research claim, at a lab or at a company publishing its own, has two checks per result. First, a specialist in the field has read it before the press has. Second, its provenance is written down, covering the prior papers it leans on and any user conversations that fed it. With both done, the announcement is a report of work. Thom's complaint is about missing provenance, and in that case the math can be correct while who gets the credit is in dispute [12][13]. A documented result with no specialist check gets read as a press claim until the field catches up. A result with neither, multiplied by scores of results arriving at once, is the case researchers say they dread [7].
What to watch
- Whether OpenAI publishes a schedule and per-result provenance for the scores of results from its unreleased model, and whether the panel signs off on them.
- Whether Anthropic or other labs commit to following the panel's advice, since its brief extends beyond OpenAI.
- Whether OpenAI answers Thom's training-data allegation with a disclosure of how user conversations with ChatGPT feed its math work.