Published · 5h agoBuild2 min read
No accuracy number here: Krishnan's councils kept 22% to 37% of one-model ideas
The figure circulating as four-agent group accuracy has no source in this work. What Krishnan measured is idea survival, on sixteen prompts that have no right answer.
Written for builders.See today for builders

What happened
- Krishnan ran sixteen open-ended prompts for the experiment: eight strategy problems and eight writing tasks.
- The council designs tested were: taking each model's answer and giving them to a fourth model to write the final version; an llm-council with peer review and a chairperson summary; and a 'best answer' picker that makes a direct pick.
- In the final runs the blended council kept only about a quarter of the good ideas that appeared in just one model's answer, so roughly three quarters of ideas two blind judges rated useful, non-obvious and worth keeping did not reach the final answer.
- Rare single-model ideas survived the peer-review council at about the same rate as plain blending: 24% versus 22%.
- In the peer-review council, an idea raised by several models was kept about a third of the time, against a quarter when only one model raised it.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
The band this experiment actually produces is 22 to 37 percent [13], and it counts idea survival, not correctness. The unit is a card: each answer was broken into small items using Sonnet, a mechanism, an observation, a metric, a failure mode, then cards meaning the same thing were clustered and labelled single-model if they appeared in one solo answer, shared if they appeared in more [9]. Two judges rated the solo-derived clusters blind to which model wrote them and to whether any council kept them [9]. Useful, non-obvious, worth keeping is the whole rubric [9]. There is no key, so there is no chance line to sit above or below [16].
The consensus tilt is where the arithmetic gets slippery. Krishnan puts peer review's preference for shared ideas at an 11 point uplift and calls it a 50% relative lift [7]. Eleven over twenty-two is fifty percent; eleven over twenty-four is nearer forty-six [14], and both single-model figures appear in his own write-up [4]. He also says the denominator for shared ideas is small [8], which puts the load on the thinnest count in the study.
On the four-agent framing: the only configuration with four models in it is the blend, three answers handed to a fourth model to write the final version [2]. That fourth model is a bottleneck, not a debater. What it reliably improved was prose, which Krishnan describes as calmer, more complete, less jagged [10]. The losses read the other way: scent cartridges salvaged from retail and traded as status symbols in a squatted mall, or the argument that logged-but-deprioritised risks are worse than unknown ones because they manufacture a false sense of control [17].
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
Krishnan ran sixteen open-ended prompts for the experiment: eight strategy problems and eight writing tasks.
ReportedView cited source - [2]
The council designs tested were: taking each model's answer and giving them to a fourth model to write the final version; an llm-council with peer review and a chairperson summary; and a 'best answer' picker that makes a direct pick.
ReportedView cited source - [3]
In the final runs the blended council kept only about a quarter of the good ideas that appeared in just one model's answer, so roughly three quarters of ideas two blind judges rated useful, non-obvious and worth keeping did not reach the final answer.
ReportedView cited source - [4]
Rare single-model ideas survived the peer-review council at about the same rate as plain blending: 24% versus 22%.
ReportedView cited source - [5]
In the peer-review council, an idea raised by several models was kept about a third of the time, against a quarter when only one model raised it.
ReportedView cited source - [6]
The selector surfaced about 37% of all good single-model ideas and 24% of the multiple-model ideas, because it picks one full answer and discards the others.
ReportedView cited source
Sources & coverage · 2 publishers
The reporting this story was synthesized from, earliest first. Every link goes to the original.
- substack.com5h agoLLM councils show groupthink - by Rohit Krishnan
- exponentialview.co5h ago🔮 Is AI immune to groupthink?

