Product1 distinct publisher2 min readUpdated
Rohit Krishnan cut the answers to 16 prompts into idea cards and rated them blind. The merge step, not the model choice, is where the rare material goes.
The Product Desk · Product desk

Compiled by The Product DeskSomething wrong?How this is made
The topology that lost the least is the one that read the least. Krishnan's straight picker, which chooses one whole answer and throws the rest away, still surfaced about 37 percent of the good ideas that came from only one model [8], roughly 12 points and half again more than the aggregator that read every answer and wrote a fresh final draft [1]. The loss is happening in synthesis, not in selection.
The picker also inverts the ranking. Because it keeps one voice and discards the overlap by construction, single-model ideas beat shared ones inside it, 37 percent to 24 [8], while the peer-review council does the reverse [3]. That peer-review shape is the LLM Council pattern Karpathy proposed as a way for several models to work with each other toward a better answer [2]. Krishnan puts its consensus premium at 11 points, or a 50 percent relative lift [7]. The mechanism is worth being precise about: peer review did not make rare ideas scarcer than plain blending did, 24 percent against 22 [9]. It promoted agreement instead. A committee that upgrades whatever two models both said is a consensus detector, which is a different instrument from the one implied by "use model diversity to get better responses" [1].
What drops out is legible in the misses. A field report noticing that salvaged retail scent cartridges had become status symbols in a squatted mall. An incident report arguing that logged-but-deprioritized risks are more dangerous than unknown ones, because they manufacture a false sense of control. A recovery plan that asks users to re-confirm suspect fields at their next login, quietly crowdsourcing the fix from the one authoritative source [10]. Each is the kind of specific that a second model is bought for, and each is what a chairperson writing a clean summary has no reason to carry. The summaries do read better, calmer and less jagged [12], which is the trade sitting in plain sight.
The caveats are the author's own. Sixteen prompts is a small run [3], the shared-idea denominator is smaller still [11], and the whole card-and-cluster method exists because Krishnan wanted a comparison he could run without human raters [4]. Take it as one experiment rather than a benchmark. The ceiling is still the finding: the friendliest architecture in the set dropped about 63 percent of the ideas two blind judges called useful and worth keeping [2]. If the reason for running four models is that they disagree, most of that disagreement is being paid for and then handed back at the merge.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
One way to get the best out of LLMs is model diversity: the models are not all the same, so using their unique natures can produce better responses.
Karpathy came up with LLM Council as a way to get multiple models to work with each other and produce a better answer.
Krishnan ran sixteen open-ended prompts: eight strategy problems and eight writing tasks.
Method: each answer was broken into small cards using Sonnet (a mechanism, observation, metric, failure mode or other detail); cards meaning the same thing were clustered; a cluster in one solo answer counted as a single-model idea and in more than one as shared; two judges scored the solo-derived clusters without knowing which model produced them or whether a council kept them. Krishnan describes it as the cleanest test he could find without doing human rating.
In the final runs the blended council kept only about a quarter of the good ideas that appeared in just one model's answer; roughly three quarters of ideas two blind judges rated useful, non-obvious and worth keeping did not make it into the final answer.
In the peer-review council, an idea raised by several models was kept about a third of the time, while an idea raised by only one model was kept about a quarter of the time.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One transparent but small self-run experiment
The method is spelled out end to end — card decomposition, clustering, single-model versus shared labelling, two judges blind to model identity and to council retention — and the results are reported as specific coverage percentages per topology. But it is 16 prompts, one author, one run, with card extraction performed by an LLM (Sonnet) that is itself unvalidated here, no inter-judge agreement statistics, no confidence intervals, and an explicitly small shared-idea denominator. Nothing is externally replicated.
No deployment or usage evidence supplied
The sources contain one researcher's experiment plus a passing reference to Karpathy's llm-council and prior MarketBench work. There are no deployment counts, usage disclosures, pricing data or production reports for council-style aggregation, so adoption cannot be scored without inventing facts.
Mildly overstated precision, deflationary direction
The framing runs against hype rather than with it — the piece argues councils are lossy and not cheaper — so the gap is small. It is positive rather than zero because precise-sounding figures ('11% uplift', 'a 50% relative lift!', '37%') are drawn from 16 prompts and an admittedly small shared-idea denominator, and the headline generalisation about councils rests entirely on this one unreplicated run.
Independent practitioner blog, attention incentive only
The material is a self-published Substack essay by an individual researcher with no vendor, model provider or product being sold in the supplied text, and it undercuts a popular multi-agent narrative rather than promoting a tool. The residual incentive is the attention value of a counterintuitive result and the absence of any independent review before publication.
Directionally credible, numerically soft
The direction — aggregation discards a majority of rare high-value ideas, and topology decides which ideas survive — is consistent across all three tested configurations and is echoed by the cited group-decision literature on biased sampling of shared information. The specific percentages deserve much less confidence: single source, single run, 16 prompts, LLM-assisted card extraction, small shared denominator, no adoption evidence.
build
The judge went synthetic first, which tells you which part of your pipeline is next1 distinct publisher
leadership
The year's most useful AI tool at one desk was a folder full of Markdown1 distinct publisher
build
A Retention Policy for Agent Memory: Flag Unused Skills at 30 Days, Archive at 901 distinct publisher
build
Query-aware compression: AWS bets a second model call is cheaper than a fat RAG prompt1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 23, 2026