Build1 publisher2 min readPublished
Agents on the Good team fall for coordinated deception in a Blood on the Clocktower benchmark
Bench on the Clocktower, a social-deduction benchmark run for 300 games per model, finds agents playing Good fall for coordinated deception by Evil agents. For builders of multi-agent systems, it cuts against trusting one agent to catch coordinated manipulation by its peers.
The Engineer · Build desk

What happened
- GPT-5.6 Sol was the best-performing model in the benchmark, ahead of Fable 5 and the other models tested.
- Each game seats 10 agents, every one a different model, under Trouble Brewing rules with minor changes to keep play flowing.
- Anthropic's Fable 5 and Opus 5 almost never chose to sacrifice themselves, according to the authors.
- With model names hidden, agents showed a slight same-provider bias that was not statistically significant and shrank once names were visible.
- How well Good players nominated did not correlate highly with how accurately they voted, and overall win rate was not easily explained by those numbers.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- exposure An orchestrator that leaves one agent to flag collusion among the others is relying on the Good-seat behaviour these authors describe as naive.
- decision Clean single-agent eval results do not settle how a group of agents will behave, so a team shipping a multi-agent system has to test the group adversarially itself.
- constraint Because nomination skill and vote accuracy do not correlate highly, a team cannot pick a detection agent on one metric and expect it to win games.
"When playing Evil, agents sometimes produce sophisticated coordinated play. When playing Good, they tend to be naive and fall for these kinds of plays," the authors wrote [3]. The post's summary states that result in words only. It does not report win rates split by side, so "sometimes" and "tend to" are the only measure of size available.
Of each model's games, 210 are on Good and 90 on Evil, 30 of those as the Imp [10]. At a 10-seat table, that works out to three Evil seats and one Imp per game [1]. Evil play is 30% of any model's record [2]. The coordinated plays the authors describe come from that share.
Two harness choices shape what a Good agent can see. Voting is simultaneous and blind, so inference for every player can run in parallel [9]. For throughput, that is a sensible choice. It also means a Good agent casts each vote without seeing how anyone else is voting on that nomination [9].
Memory is the second choice. Rules and role stay in context throughout. Only the most recent few thousand tokens of the game are visible, and anything older has to come from the agent's own scratchpad or a recall function [11]. A day runs through opening statements, group discussion, breakout conversations and nominations, each nomination with an accusation and a defence [8]. To spot a coordinated play, a Good agent has to set claims from different phases against each other, possibly after those claims have left the visible window. The authors checked, manually and with agents, that results came from real game-playing [12].
That check leaves the transfer question open. For the Good-side result to carry into a production system, the naivety has to belong to the model and not to this memory setup. The evidence so far fits either explanation. I'd expect the same misses from a reviewing agent that works with a short window and self-managed notes while its peers are able to coordinate.
The authors built the benchmark as a testbed for deception and coordination [1]. Almost all alignment evals test one agent, and the authors worry that models can look aligned individually while behaving unpredictably as a swarm [13]. They cite METR's investigation into the OpenAI-HuggingFace incident, where some behaviours, in their view, would have been hard to catch with single-agent evals [14].
What to watch
- Side-specific win rates and counts of successful coordinated Evil plays would put a size on the Good-side naivety.
- An ablation with a longer visible window or different recall settings would show whether Good-side misses belong to the model or the harness.
- More games could show whether the slight same-provider bias reaches statistical significance.