Skip to content

Build1 publisher3 min readPublished

Multi-agent LLM gains largely vanish once the thinking-token budget is held constant

An arXiv preprint argues single agents are more information-efficient under a fixed reasoning budget, and reports they match or beat orchestrated agents on multi-hop reasoning.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • The paper finds that single-agent systems consistently match or outperform multi-agent systems on multi-hop reasoning tasks when reasoning tokens are held constant.
  • The authors conclude that for multi-hop reasoning tasks, many reported advantages of multi-agent systems are better explained by unaccounted computation and context effects than by inherent architectural benefits.
  • The paper defines thinking token budgets as the total number of tokens used for intermediate reasoning, excluding prompts and final answers.
  • Multi-agent systems typically consume more tokens through longer reasoning traces or multiple agent interactions, making it unclear whether their gains arise from architectural advantages or simply from increased compute.
  • The paper presents an information-theoretic argument grounded in the Data Processing Inequality: under a fixed reasoning-token budget and with perfect context utilization, single-agent systems are more information-efficient, because multi-agent decompositions introduce additional communication bottlenecks that can lead to information loss.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A preprint posted to arXiv reports that when reasoning tokens are held constant, single-agent systems consistently match or outperform multi-agent systems on multi-hop reasoning tasks [1]. The paper's framing is the useful part for anyone shipping an orchestrator: reported multi-agent gains are, in its account, better explained by unaccounted computation and context effects than by inherent architectural benefit [2].

The unit of control matters. The authors define a thinking token budget as the total number of tokens used for intermediate reasoning, excluding prompts and final answers [3]. Multi-agent systems typically consume more of those tokens, through longer traces or multiple agent interactions, which leaves it unclear whether their gains come from architecture or simply from more compute [4]. That is the confound: the comparison most teams run internally is not architecture versus architecture, it is architecture versus a larger bill.

The theoretical claim is an information-theoretic one, grounded in the Data Processing Inequality: under a fixed reasoning-token budget and with perfect context utilization, a single agent is more information-efficient, because decomposing reasoning across agents introduces communication bottlenecks that can lose information [5]. Multi-agent systems decompose reasoning across agents that operate over partial contexts and communicate via generated text, while a single agent reasons within one unified context [6]. The same argument produces the condition under which orchestration earns its keep: when a single agent's effective context utilization is degraded, for example by long or noisy contexts, or when the multi-agent system is quietly spending more compute [7]. That is the test to apply before adding a planner. Not "does the swarm score higher", but "is my single agent demonstrably failing to use the context I am giving it".

The empirical work compares single-agent setups against multiple multi-agent architectures under matched reasoning-token budgets across three model families: Qwen3, DeepSeek-R1-Distill-Llama, and Gemini 2.5 [8]. The taxonomy of systems the paper is arguing about is broad, covering planners, role-playing systems, debate frameworks, and tool-specialized swarms [9]. The authors also note that earlier budget-aware studies have already found many such strategies underperforming strong single-agent baselines once computation is normalized [10].

Two diagnostic findings deserve attention independently of the headline result. First, the authors identify significant artifacts in API-based budget control, particularly in Gemini 2.5, which distort effective computation [11]. If you are enforcing a reasoning budget through a vendor parameter, your normalization may not be doing what you think it is doing. Second, they report benchmark vulnerabilities exposed through paraphrasing, and note that both the budget-control and benchmark artifacts can inflate apparent multi-agent gains [12]. They also report systematic differences in failure modes across architectures [13].

The honest caveat: the material available here is the abstract and the opening of the introduction, which report no accuracy figures, name no specific benchmarks, and give no absolute token budgets [14]. The direction of the result is stated plainly; its size is not.

What to watch: whether the API budget-control artifact reproduces outside Gemini 2.5, since that determines whether token-normalized evaluation is even reproducible on hosted models [11]. Watch also whether vendors publishing multi-agent benchmark wins begin reporting reasoning-token counts alongside scores [4], and whether the paraphrase sensitivity holds up on the multi-hop benchmarks teams currently use for internal gating [12].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories