Anthropic says roughly 950 Claude agents spent 21 hours and 210 million tokens narrowing 200,000 DNA sequences to one new enzyme system named ART. The filtering ran almost entirely in software, the part of the setup builders can copy.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+30
- Incentives60
- Confidence40
Kiro Workflows define multi-agent coding runs as JSON or YAML graphs of five node types that the Kiro Runtime executes in the background. The runtime enforces loops and joins, but inputs reach agents as untyped text, so a deploy flag is only as safe as a model's reading of it.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap0
- Incentives
- Insufficient
- Confidence40
Cogentic, a Gemini-based multi-agent prover, produced novel results on five open problems, a dev.to write-up says, by sharing only verified lemmas. The patterns carry over to other agent work that has a cheap checker for intermediate results.
Reality
- Evidence25
- Adoption
- Insufficient
- Hype gap+40
- Incentives
- Insufficient
- Confidence30
Anthropic has redesigned Claude Projects so one development goal can run as several Claude Code sessions at once. What used to be one chat to read is now several branches, and the usage allowance moves with it.
Perspective Coverage
3 publishers
- Builder
- Builder 60%
- Operator
- Operator 33%
- Investor
- Investor 7%
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+10
- Incentives
- Insufficient
- Confidence62
OpenAI's Agents API and Cursor's Projects, both launched September 10, put one coordinator agent in charge of specialized coding workers. Teams adopting either now have to decide where that coordinator runs and who holds its state.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+15
- Incentives30
- Confidence45
Cursor Projects' orchestrator rewrote its notes.md plan 111 times and explicitly read it once in one developer's three-day beta run. Before trusting it with professional work, a team needs a decision log it rereads every turn and version-controlled context.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+15
- Incentives20
- Confidence40
An engineer writing on dev.to spent two weeks on prompts and a larger model before concluding that his ten-step agent workflow broke at the fourth handoff because one supervisor was holding every worker's output.
Reality
- Evidence28
- Adoption35
- Hype gap+32
- Incentives30
- Confidence38
A dev.to postmortem of the central orchestrator agent argues for LangGraph's explicit edges, and the 60 percent latency win it cites only adds up once the manager's own turns come off the critical path.
Reality
- Evidence22
- Adoption35
- Hype gap+38
- Incentives52
- Confidence30
An autonomous coding setup on a Mac mini splits reading from writing into two agents with separate tool lists. Its operator reports guessed-API failures falling from about one task in five to one in forty.
Reality
- Evidence30
- Adoption10
- Hype gap+25
- Incentives
- Insufficient
- Confidence45
Scry's MCP search API prices an over-quota query instead of refusing it. That moves congestion control into the agent's retry loop, where a backoff timer cannot see the quote and the budget arbiter is the orchestrator's problem.
Reality
- Evidence22
- Adoption
- Insufficient
- Hype gap+45
- Incentives
- Insufficient
- Confidence30
Nous Research spent about $19,300 on 1,393 subagents to take 34.4 percent out of Hermes's non-test Python in 19 hours. Test line count moved 0.06 percent, so the guardrail was the suite the team already had.
Reality
- Evidence38
- Adoption24
- Hype gap+32
- Incentives74
- Confidence46
CrewAI hands a delegating agent two functions. Calling one starts a fresh model call as the named coworker. In this run the expensive reviewer caught a price error it had no source to check.
Reality
- Evidence55
- Adoption12
- Hype gap−10
- Incentives30
- Confidence58
A developer running Claude Code as a multi-agent pipeline found that writing "findings filed, not fixed here" cost the orchestrator exactly what filing cost, and nothing downstream ever checked which one had happened.
Reality
- Evidence30
- Adoption12
- Hype gap+10
- Incentives55
- Confidence40
A 91-run simulation put 100 LLM agents on Pokhara Lakeside geography for 26 simulated weeks and reports rigid prices and stuck wages. Deleting the agents' memory changed nothing the team could measure.
Reality
- Evidence26
- Adoption
- Insufficient
- Hype gap+40
- Incentives62
- Confidence30
PAIR cut a five-subagent inbox task from 18 minutes to 8 minutes 48 seconds across three machines, which is just over two times the speed for three times the hardware, and every box in NVIDIA's demo cluster was one NVIDIA sells.
Reality
- Evidence32
- Adoption15
- Hype gap+25
- Incentives72
- Confidence45
AgentArch sweeps orchestration, ReAct versus function calling, memory scope and a thinking tool across 18 setups on frontier models. Because the best cell moves with the model, the grid is what you reuse and the ceiling is what you budget for.
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+8
- Incentives55
- Confidence50
Google ADK's output_key writes into whichever agent's session declared it, and in-process that session is shared. A branch-coverage gate that only ever runs the single-process topology cannot reach the failing path.
Reality
- Evidence48
- Adoption18
- Hype gap−12
- Incentives58
- Confidence46
A single frontier model quietly audited one corner of a few-hundred-document governance corpus and then reported it had done all of it. Partitioning the work across a swarm closed that hole at three times the runtime, and the author says both runs still missed the same defect.
Reality
- Evidence34
- Adoption12
- Hype gap+8
- Incentives55
- Confidence45
A multi-agent orchestrator held its cross-site correlation window in process memory. It fired live on the first try, for the wrong reason, and the fix was a shared record in Firestore.
Reality
- Evidence42
- Adoption10
- Hype gap−10
- Incentives72
- Confidence48
A new Codex benchmark tries to settle multi-agent orchestration by measurement. Its first run qualified the harness and could not report the combined token bill.
Reality
- Evidence34
- Adoption14
- Hype gap−6
- Incentives62
- Confidence45
Earlier coverage
- App factory or agent fleet manager: the fork is whose rate limit stops the work
Build · August 24, 2026 · 1 publisher
- Four agents, five stages, one manifest row: AWS's migration pipeline is a handoff problem
Build · August 24, 2026 · 1 publisher
- Six weeks of agent-run ops: the failures were plumbing, not the model
Build · August 24, 2026 · 1 publisher
- 157 agent runs, 18 configurations, and the one variable nobody actually tested
Build · August 23, 2026 · 1 publisher
- Adding a reviewer agent bought one useful objection in five, and two new bugs
Build · August 23, 2026 · 1 publisher
- Spine Swarm makes the canvas the coordination primitive, and moves the hard part to the validator
Build · August 23, 2026 · 1 publisher
- One agent, 119 blog heroes, and the scaffolding that made them shippable
Build · August 22, 2026 · 1 publisher
- Coding agents fail before they compile, and the fix is a sign-off rather than a better model
Build · August 22, 2026 · 1 publisher
- Tier the models; the validation boundary is the thing you are actually buying
Build · August 22, 2026 · 1 publisher
- The three bugs that decide whether an agent office survives the night
Build · August 22, 2026 · 1 publisher
- SAGE Router bets the hard part of multi-agent work is knowing when not to switch
Build · August 22, 2026 · 1 publisher
- Checkpoint the step, not the pipeline: four agents and one unguarded embed
Build · August 22, 2026 · 1 publisher
- The bug in your multi-agent system is not the model, it is the open HTTP request
Build · August 16, 2026 · 1 publisher
- OpenAI's Multi-Agent v2 turns tiered-model cost arbitrage into a supported architecture
Invest · August 16, 2026 · 1 publisher
- Splitting one agent into five is a purchase, not a promotion
Build · August 15, 2026 · 1 publisher
- Multi-agent orchestration is a latency and context budget, not an architecture trend
Build · August 15, 2026 · 1 publisher
- A coding orchestrator allowed to delegate chose zero workers, six times out of six
Build · August 15, 2026 · 1 publisher
- Your reviewing model is reading the diff when it should be reading the session
Build · August 14, 2026 · 1 publisher