Skip to content

Build1 publisher2 min readPublished

Noam Brown credits under 10% of OpenAI's Navier-Stokes swarm result to agent cooperation

A LessWrong post argues organized swarms could turn parallel test-time compute superlinear. Its two exhibits are unreleased OpenAI runs, and the only estimate on record gives coordination less than a tenth of the credit.

The Engineer · Build desk

Photograph accompanying Noam Brown credits under 10% of OpenAI's Navier-Stokes swarm result to agent cooperation
Photo: substack.com

What happened

  • A LessWrong post argues that swarm organization could move parallel test-time compute from a sublinear to a superlinear exponent, so adding parallel agents starts buying increasing returns instead of diminishing ones.
  • Its evidence is two runs inside OpenAI: the 700-agent swarm behind the Hugging Face attack, and the 10,000-agent swarm that solved the Navier-Stokes Millennium Prize problem.
  • Both feats were largely performed by unreleased internal models, and the post says how much swarm organization contributed to either outcome is not known.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • contradiction Superlinearity is offered as the reason to fund coordination. But the only quantitative estimate in the post assigns cooperation under a tenth of its flagship result, so the case for the investment rests on analogy to human organizations.
  • constraint Both demonstrations ran on unreleased internal models, so no outside team can hold the model fixed and vary the coordination layer. That experiment would settle whether returns are sublinear or superlinear.
  • decision Brown's account of the message board suggests cooperation is a property the model was trained into. If so, budget belongs in models trained to cooperate before it goes into messaging and shared-state plumbing around a released one.
  • precedent A 10^4 headcount is now the reference point for a swarm demo, and the next round of announcements will be graded on agents launched unless someone publishes capability per agent.

Superlinear returns to parallel agents would show up as a curve. Capability on one axis, agent count on the other, with the model held fixed and the coordination layer varied. The post has no such curve. On whether returns are already superlinear in domains like cybersecurity and mathematics, it says "this remains to be rigorously measured" [9], and it says the contribution of swarm organization to the two OpenAI outcomes is unknown [6].

That leaves the two runs as the evidence. For the Hugging Face attack, the post reports that Noam Brown of OpenAI believes the swarm's use of an internal message board came out of multi-agent training, and that it was not a proper example of it [7]. For Navier-Stokes, the post reports Brown's estimate that under 10% of the result was due to multi-agent cooperation [8]. Take the estimate at face value and more than 90% of the strongest exhibit for organized swarms came from something other than agents cooperating [10]. Two runs is a thin population for a scaling claim.

The second argument is growth in swarm size. Claude Code shipped subagents in July 2025, used mostly in tens [11]. Swarms in the hundreds were reported more regularly through 2026 [12]. OpenAI ran an effective swarm of 1,000 in July 2026 and 10,000 in September [13]. That is three orders of magnitude in fourteen months, roughly one order every 4.7 months [14]. The series counts agents launched. Inside that count, a bigger pool of parallel samples and better cooperation between them look the same.

The rest is analogy and ceiling. 400,000 coordinating humans landed people on the moon [15]. Agent transaction costs should run below human ones, the post argues, particularly when siblings share a model and harness and have tools to verify each other's identity [17], and on a Coasean reading that would let a corporation of 100 million agents stay effective [16]. Agents can also share skills and memories directly and work without handoffs [18].

For any of it to transfer to a stack you can buy, two things have to hold. The model has to be trained to cooperate, since Brown's account attributes the message board to multi-agent training rather than to the harness [7]. The gain then has to survive on released weights, and both OpenAI runs were largely performed by unreleased internal models [5]. Without both, a message bus in front of a model that was never trained to use one buys parallel sampling. The post itself treats that as sublinear [21].

What to watch

  • A published curve of capability against agent count on released weights, with the coordination layer varied and the model held fixed.
  • Whether OpenAI publishes the models, harness, or message board logs behind the 700-agent and 10,000-agent runs.
  • Whether Brown revises the under-10% attribution as OpenAI describes its multi-agent training in more detail.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories