Skip to content

Build1 publisher3 min readPublished

Wiring, not headcount: same agent task swung from 70% worse to 81% better on topology alone

A summary of work attributed to Google Research and MIT reports 180 controlled runs where multi-agent setups averaged +0.2% against one agent. The spread came from how agents were connected.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Wiring, not headcount: same agent task swung from 70% worse to 81% better on topology alone
Generated illustration

What happened

  • A dev.to article headlines the claim that with the model and tools unchanged, changing only the connection map between agents flipped results from 70% worse to 80% better.
  • The dev.to article is dated 15 August 2026, bylined Nokka, and states it was written by an AI (DeepSeek V4 Pro) via Hermes Agent under human control and quality review by Nokka.
  • An X post by Miraqle (@0xMiraqle), reported at 101.9K views, states: a Google Research and MIT team ran the same agent jobs hundreds of different ways with identical prompts, identical tools and identical compute budget, changing only how the agents were wired to each other; the same work swung from 70% worse than a single agent to 80% better and averaged out to basically zero.
  • The research is named "Scaling Multi-Agent Systems", attributed to a Google Research and MIT team, testing 180 configurations with three models (GPT, Gemini, Claude) and five different agent-connection architectures.
  • The Decoder is quoted as summarising: multi-agent systems swung wildly in performance depending on the task, from an 81 percent boost to a 70 percent drop, across 180 controlled experiments spanning five architecture types and three model families.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A summary of research attributed to Google Research and MIT says that when prompts, tools and compute budget were held identical and only the connection graph between agents was changed, the same job ran anywhere from 70 percent worse than a single agent to about 80 percent better, and averaged out to basically zero [3]. Read narrowly, that kills "add more agents" as an improvement strategy: the variable being tuned is not headcount, it is topology [1].

The reported setup was 180 configurations across five agent-connection architectures and three model families, given as GPT, Gemini and Claude [4]. The Decoder's account of the same work describes 180 controlled experiments across five architecture types and three model families, with performance swinging from an 81 percent boost to a 70 percent drop [5].

The split is along task shape. On parallelizable work, with financial statement analysis as the example, centralized coordination is reported to have improved results by 80.9 percent against a single agent [6]. On sequential reasoning, with planning as the example, every multi-agent variant tested degraded performance by 39 to 70 percent, according to a summary from evoailabs [7]. That is roughly a 151 point range attributable to wiring [1], around which the overall mean sits at +0.2 percent [8], about a tenth of a percent of the spread [4]. An average that small is not evidence that multi-agent is fine; it is evidence that the wins and the failures are the same size and you are picking one blind.

The stated mechanism is error propagation, not coordination overhead. Crews with no correction step are reported to have multiplied their own errors up to 17 times the solo rate, while supervised crews held it near 4 times [9], which is roughly a fourfold reduction from putting one checker above a fan-out [2]. The companion rule is to stop agents reading each other's drafts, so a wrong step reaches the supervisor instead of infecting four other bots [10]. Each worker gets one job and one output [14].

Two operating rules are worth more than the percentages. First, keep one lone agent running as the control, because it is the only number that tells you the crew is earning its calls [12]. Second, re-run the solo-versus-crew test after every model upgrade, since a smarter base model quietly makes the crew stop paying [11]. The sharpest claim in the writeup is a threshold: if the single agent already clears about 45 percent success rate, building a crew around it returns zero or negative [13].

Provenance deserves flagging. The article carrying these numbers was published on dev.to on 15 August 2026 under the byline Nokka and states it was written by an AI, DeepSeek V4 Pro, via Hermes Agent under human control and quality review [2]. Its numbers are routed through an X post by Miraqle that had 101.9K views [3], The Decoder [5] and evoailabs [7], and the piece does not link a primary paper [15]. The two headline figures already disagree by a point, 80 percent versus 81 percent [3]. Treat the shape as the finding and the decimals as unverified.

Watch for the paper itself, titled "Scaling Multi-Agent Systems" in this account [4], and specifically for whether the 39 to 70 percent sequential penalty [7] survives with reasoning-heavy base models, since the same writeup predicts stronger models erode crew value [11]. Also watch whether any orchestration vendor publishes a single-agent control alongside its benchmarks [12].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories