Skip to content

Build1 publisher3 min readPublished

Splitting one agent into five is a purchase, not a promotion

A dev.to head-to-head scores single-agent and multi-agent designs on seven production criteria. Multi-agent wins scope and security isolation, and pays for it in latency, cost and glue code.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • A founder at a workflow-automation startup had a working AI agent handling customer onboarding (one agent, a set of tools, a system prompt), then read framework marketing pages that all pushed the message that agents are meant to be teams (researcher, writer, reviewer, operator), and asked whether to split into five specialized agents.
  • The author's first question was "What is breaking today that five agents would fix?" and the honest answer was nothing.
  • The article's thesis: multi-agent systems are a tool, not a trend; most teams should run a single agent, a minority should split, and multi-agent gives you scope, isolation and security boundaries at the cost of latency, cost and complexity.
  • The author scores both architectures on a 1-5 scale where 5 is best, stating these are production scores from having built and shipped both, not feature-checkbox scores.
  • Latency: single agent 5, multi-agent 2. In the author's load tests a simple five-agent pipeline ran 4-8x slower than a single agent on the same task.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A founder at a workflow-automation startup asked an engineer whether to split a working customer-onboarding agent into five specialized ones, after every framework marketing page he read told him agents are meant to be teams [1]. The engineer's first question back was what is breaking today that five agents would fix, and the honest answer was nothing [2] - which is the whole decision in one exchange, because the published scoring that follows shows multi-agent buying you real things at a price most teams are not yet forced to pay [3].

The writeup scores both architectures 1 to 5 across seven criteria, based on the author's own production experience rather than feature checklists [4]. Single agent takes latency 5 to 2 [5], cost 4 to 2 [6], reliability 3 to 2 [9] and maintainability 5 to 2 [10]. Multi-agent takes context and tool surface 5 to 2 [8], security isolation 5 to 2 [11] and failure isolation 4 to 2 [12]. Totals: 23 to 22 [13]. That is not a verdict, it is a warning that the two columns are close and the reasons behind each number are what matter. Note also that these are one practitioner's scores, single-sourced.

The costs are the concrete part. Every handoff re-serializes context into a fresh prompt, so a five-agent pipeline ran 4 to 8 times slower than a single agent on the same task in the author's load tests [5][7]. Token cost multiplies because each agent re-reads shared context and each handoff duplicates history: a task that cost $0.03 as a single agent cost $0.11 as a five-agent crew, which the author calls roughly 4x [6]. At 100,000 tasks a month, that 8 cent delta is $8,000 [14]. "The models are cheap" holds until volume makes the multiplier the line item [6].

What you buy is scope and permissions. One context window holds only so much, and a single agent juggling a large codebase, a long conversation and 20 tools starts forgetting that tools exist and trimming its own context [8]. Splitting gives each agent a small window and a small tool set [8]. The stronger argument is security: in a single agent every tool is equally reachable, so the component that reads customer data can also write to the database, while a split lets a read-only researcher never hold write credentials and a write-capable operator never see raw PII [11]. Failure isolation follows - a failing sub-agent can be retried or swapped without restarting the task [12].

The bill nobody budgets is the orchestration layer, which the author describes as a small distributed system rather than several agents [15]. Handoff contracts are not defined by default, so they drift until the reviewer agent is parsing prose the writer formatted as a table, and you end up writing schemas for inter-agent communication [16]. Loop control is its own question: who decides when the crew is done [17]. Multi-agent also adds failure modes single agents do not have, including lost information at each boundary and crews arguing in circles [9]. Debugging means five prompts, five context snapshots and the orchestration logic that picks handoffs [10].

Watch for the trigger conditions rather than the architecture diagram. Split when your agent is visibly dropping tools or self-trimming context [8], when read and write paths must not share credentials [11], or when a sub-task needs to fail and retry alone [12]. Until one of those is true, the 4x token multiplier and the 4 to 8x latency penalty are being spent on nothing [5][6].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories