Skip to content

Build1 publisher3 min readPublished

Werewolf agents read the harness's own turn order as evidence of guilt

Four Fable 5 agents playing Werewolf through Hyperagent kept citing silences that the turn order had imposed. Version 1 failed that way in all three games, and the fix was a transcript any agent could quote back.

The Engineer · Build desk

Illustration accompanying Werewolf agents read the harness's own turn order as evidence of guilt

What happened

  • Four Fable 5 agents were given hidden roles to play Werewolf through Hyperagent, and the opening speaker accused Ptolemy of a telling silence when Ptolemy had simply not been handed a turn yet.
  • Version 3 added a canonical public record and explicit private state, so a claim could be checked against a transcript held outside any single agent's memory.
  • Late in version 3, Ptolemy said he conceded in full after Finch quoted the transcript showing Finch had named Ptolemy before Inez spoke at all.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint With no state changes between speeches, extra discussion rounds in version 1 could only add interpretations of the same empty record. Turn count was not a source of grounding.
  • decision Harness builders have to decide which facts are labelled as the operator's, because a randomized speaking order arrived unlabelled and was read as a player's deliberate choice.
  • exposure Any evaluation that grades a multi-agent run on how the transcript reads will pass the ungrounded version, since the false claims are written in competent social-deduction language.
  • contradiction The failure is counted across three of three games while the fix rests on a single exchange in a version that also changed headcount and format, so the transcript cannot yet be credited on its own.

Version 1 gave each agent a role, a turn, and a prompt asking what the table's behavior meant. That was the environment. Four players, one werewolf, three public discussion rounds, no eliminations, no night kills, one vote at the end [2]. Nothing in that loop changes between speeches except who has spoken. Sable spoke first, so at the moment of the accusation the public record held zero prior speeches [21].

The accusation was this: "Ptolemy. Hasn't said a word yet and that silence is doing a lot of work." [3] Ptolemy had not spoken because it was not his turn [4]. Bosch, going second, ran the same move on Wren: "Wren has said nothing, which is precisely what a careful operator does when the opening move belongs to someone else." [5] Wren had not received a turn either [6]. The second game opened with Bosch naming Ptolemy as the one to watch, before Ptolemy had spoken [7].

In the third game Ptolemy asked Wren, "How did you decide on the order in which you were going to speak today?" [8] The Game Master had randomized the order [9]. That is the plumbing failure in its clearest form. An operator-controlled variable reached the agent with nothing marking who owned it, and the agent handed it to a player as a decision.

The pattern held in all three version-1 games, with at least one agent treating behavior that had not occurred, or a condition the operator controlled, as evidence about another player [10]. Villagers did it as well as the werewolf, so bluffing does not explain it [11]. The author wrote: "Reasoning-shaped language is not the same thing as reasoning over the environment." [12]

The change that bears on grounding is a canonical public record plus explicit private state [15]. Votes changed who remained alive. Night actions changed the next day's state, eliminated roles were revealed, the Seer received private information, and the werewolves coordinated privately [16]. Claims could then be checked against a transcript that existed outside any single agent's memory [15]. Late in version 3, Ptolemy claimed Finch's position had followed Inez's lead [17]. Finch said: "The transcript has me naming Ptolemy in my own speech before Inez ever opened her mouth." [18] Ptolemy said: "Yes, I concede in full... the transcript has Finch's 'leans Ptolemy' before Inez ever spoke." [19]

How much weight that correction carries is limited by the design. Version 2 already added seven recurring characters, two werewolves, a Seer, daily elimination votes and night kills [13], and version 3 went to nine characters with social missions, private confessionals, an open floor and separate mystery and omniscient audience cuts [14], so the transcript is not isolated as the cause. The failure is documented in three of three games; the correction is one exchange [22]. The writeup reports no per-game count of ungrounded claims for the later versions and no controlled comparison [23]. Metered usage was covered by Hyperagent credits and the run figures are platform-reported [20], so there is no cost number here to size the fix against.

This result transfers only where the agents are asked to interpret a shared history, that history sits outside each agent's context so it can be quoted, and the harness marks which facts belong to the operator. Where those hold, a bad claim gets contradicted with a citation, and the citation is duller reading than the invented motive it replaces. Where they do not, the fluent run costs the same as the grounded one and scores the same on any rubric that reads the transcript.

What to watch

  • A per-game count of ungrounded claims in versions 2 and 3, which would show whether the canonical record lowers the rate or only makes bad claims correctable.
  • An ablation that adds the shared transcript to the four-player version-1 format while holding headcount and rules fixed.
  • Token or credit figures for the larger runs, since the writeup only reports platform run figures covered by Hyperagent credits.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories