Skip to content

Invest1 publisher3 min readPublished

Google's 180-config sweep: extra agents cut sequential-task scores by 39 to 70%

A controlled evaluation across four benchmarks found centralized coordination lifted financial reasoning 80.9%, while every multi-agent variant tested made strict sequential planning worse.

The Investor · Invest desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Google's 180-config sweep: extra agents cut sequential-task scores by 39 to 70%
Generated illustration

What happened

  • A Google Research blog post dated January 28, 2026, summarising the paper "Towards a Science of Scaling Agent Systems", reports a controlled evaluation of 180 agent configurations deriving quantitative scaling principles for AI agent systems, finding multi-agent coordination dramatically improves performance on parallelizable tasks but degrades it on sequential ones, and introduces a predictive model that identifies the optimal architecture for 87% of unseen tasks.
  • The blog post is authored by Yubin Kim, Research Intern, and Xin Liu, Senior Research Scientist, Google Research.
  • The team evaluated five canonical architectures: one single-agent system (SAS) and four multi-agent variants (independent, centralized, decentralized, and hybrid).
  • The architectures were tested across four benchmarks: Finance-Agent (financial reasoning), BrowseComp-Plus (web navigation), PlanCraft (planning) and Workbench (tool use).
  • The architectures were evaluated across three leading model families: OpenAI GPT, Google Gemini and Anthropic Claude.

Compiled by The InvestorSomething wrong?How this is made

Why it matters

Google Research published a blog post on January 28, 2026 summarising a paper, "Towards a Science of Scaling Agent Systems," based on a controlled evaluation of 180 agent configurations [1]. The headline result cuts against the prevailing build instinct: multi-agent coordination sharply improved performance on parallelizable tasks and degraded it on sequential ones [1].

The setup is worth reading before the numbers. The authors, research intern Yubin Kim and senior research scientist Xin Liu, tested five canonical architectures, one single-agent system plus four multi-agent variants labelled independent, centralized, decentralized and hybrid [2][3]. Those ran across four benchmarks: Finance-Agent for financial reasoning, BrowseComp-Plus for web navigation, PlanCraft for planning and Workbench for tool use [4], and across three model families, OpenAI GPT, Google Gemini and Anthropic Claude [5].

On Finance-Agent, where sub-problems can be split up so separate agents examine revenue trends, cost structures and market comparisons at the same time, centralized coordination improved performance by 80.9% over a single agent [6]. On PlanCraft, which demands strict sequential reasoning, every multi-agent variant the team tested degraded performance, by 39% to 70% [7]. That is a spread of roughly 151 percentage points between the best and worst reported outcomes of the same design decision [8]. The stated mechanism is not exotic: communication overhead fragmented the reasoning process and left insufficient "cognitive budget" for the task itself, according to the post [9].

The second finding is the one that should worry anyone shipping tool-heavy agents. Google describes a "tool-coordination trade-off," in which the coordination tax rises as a task requires more tools, using a coding agent with access to 16 or more tools as its example [10]. Most enterprise workflows that people are currently decomposing into agent teams look more like that coding agent than like a finance research task.

Google positions this against published work in the other direction. The post cites "More Agents Is All You Need," which reported that LLM performance scales with agent count, and collaborative scaling research finding that multi-agent collaboration often surpasses each individual through collective reasoning [11]. Its counter-argument is that the "more agents" approach hits a ceiling and can degrade results when the architecture is not matched to the properties of the task [12]. The framing rests on the observation that agents run sustained multi-step interactions in which a single error cascades through the workflow [13].

The commercial payload is a predictive model that, per Google, identifies the optimal architecture for 87% of unseen tasks [1]. By the same figure it picks wrong on 13% [14], which for a routing layer sitting in front of production traffic is a real error budget rather than a rounding detail.

Two things to watch. First, whether the selection model is released in a form operators can test, since 87% on unseen tasks drawn from four benchmarks is not the same as 87% on your backlog. Second, the missing numbers: the excerpt quantifies only Finance-Agent and PlanCraft, and shows the BrowseComp-Plus and Workbench comparisons in box plots [15], so the middle of the distribution, where most real workflows sit, is still unpriced. Until then, the cheap move is the boring one: keep sequential work on one agent and make the burden of proof fall on the team that wants to add a second.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories