Token use alone explained 80 percent of the variance on BrowseComp, and a Berkeley-led trace study found most multi-agent failures are structural, so the fan-out design pays only where subtasks are independent.
Reality
- Evidence54
- Adoption36
- Hype gap+12
- Incentives74
- Confidence55
A two-model code review produced rebuttals and a clean verdict while the raw logs showed no changed positions and no new evidence. The rebuild enforces independence by withholding each verdict until both are committed.
Reality
- Evidence34
- Adoption10
- Hype gap+22
- Incentives48
- Confidence44
The team behind a nine-agent code generator reports about 30,000 tokens and 15 to 20 Pro-tier model calls per request, and attributes its reliability to strict output schemas and context the agents cannot decline to read.
Reality
- Evidence40
- Adoption28
- Hype gap+15
- Incentives55
- Confidence42
Its agents interview the user for what the request leaves out, then build and run the modelling. The paper shows this on two enzymes, which demonstrates the pipeline rather than measuring how often it is right.
Reality
- Evidence48
- Adoption
- Insufficient
- Hype gap+30
- Incentives65
- Confidence52
An arXiv preprint argues single agents are more information-efficient under a fixed reasoning budget, and reports they match or beat orchestrated agents on multi-hop reasoning.
Reality
- Evidence52
- Adoption
- Insufficient
- Hype gap+12
- Incentives
- Insufficient
- Confidence44