Skip to content

Build1 publisher3 min readPublished

A seeded murder mystery puts a number on what swarm monitors omit

A LessWrong team fixed an eight-character murder plot before any agent spoke, then graded monitors on what they reported and what they missed. The best one fully recovered under half the rubric's facts.

The Engineer · Build desk

Photograph accompanying A seeded murder mystery puts a number on what swarm monitors omit
Photo: lesswrong.com

What happened

  • A LessWrong team fixed an eight-character murder plot set in 1930s England as ground truth, then had agents generate the interactions that monitors would have to reconstruct from the traces.
  • The strongest monitor fully recovered less than half of the rubric's facts and relationships, with omission very common and failure to connect relevant facts the other recurring error.
  • Opus 5 led the Anthropic family but reported zero reasoning tokens under the main setup, while Fable 5.1 displayed substantial reasoning under the same requested settings.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint A log-reading review cannot be the control that establishes what a swarm did, because the team running it has to budget for the part of the account no monitor reports and then decide which missing links change the conclusion.
  • decision Reasoning effort now has a measured price in recovered facts, so a team that runs monitors at low effort to hold token spend down is paying for it in omissions.
  • exposure Any harness where agents write their own summaries or conclusions into the trace gives a monitor a way to score well by reading them, and the score still looks like reconstruction.
  • contradiction Two models in the same family answered the same requested reasoning setting differently, so a cost model built on reported reasoning tokens cannot compare effort across vendors.

The design choice here is the ordering. Most reconstruction work starts from an incident and works backwards: Baig et al.'s MessageBoardAuditBench tests whether AI agents can reproduce the findings of the human investigation into the collusion.wiki incident [9]. That yields a real reference account, one case per investigation. The team's objection is practical. Accounts of that kind take substantial human investigative work, and depending on new incidents and the investigations that follow them makes cases hard to produce systematically [10]. Having an LLM write the transcript was rejected on fidelity grounds, because a fictional transcript may not faithfully represent what agents do when confronted with rules, goals and incentives [11]. So the storyline is human-authored and fixed, and the agents supply the evidence through their choices about what to do, disclose or conceal [7].

Holding the account in advance is what makes omission measurable. A grader can check both what the monitor reported and what it left out, and a reconstruction can contain accurate observations while still missing the part that links two agents [8]. Both failures showed up: omission, very commonly, and failure to connect relevant facts [1].

GPT-6 Astra at high reasoning performed best, with Astra at low reasoning second [3]. Higher effort added six percentage points of full recovery on average, with gains on all ten trajectories [4]. The gain is consistent enough to believe and small next to the gap: if each further increase bought another six points, a monitor starting under half would need roughly eight more of them to recover the whole rubric [14]. Nobody ships a dial with eight more notches.

Across vendors the reading is weaker. Opus 5 came top within the Anthropic family and reported zero reasoning tokens under the main setup, while Fable 5.1 displayed substantial reasoning under the same requested settings, which the team suspects reflects model-specific adaptive reasoning [6]. A requested effort level and the reasoning actually spent are different variables, and here they came apart inside one vendor's lineup.

Gemini's result is the one I would check for first in a deployment. It ranked worst, and its judgement correlated with that of the investigation running inside the simulation, plausibly piggy-backing on decisions made by models in the simulations [5]. A murder-mystery world contains its own detectives, so the trace contains proposed answers. If production traces carry agents' own summaries and conclusions, a monitor's score may partly measure its ability to read those.

For the under-half figure to say anything about a real incident review, the facts a reviewer needs would have to resemble the rubric's: roles, relationships and backgrounds among eight characters in a generic murder plot set in 1930s England [2]. The write-up gives the result as less than half, without a percentage or a count of rubric items. It covers ten trajectories generated from one story [4][2].

What to watch

  • Whether the team publishes per-trajectory recovery percentages and the size of the rubric graph, which would show how far under half the best monitor landed.
  • Whether reasoning effort keeps adding recovery past the first increase or flattens out after one step.
  • Whether MessageBoardAuditBench results on the collusion.wiki investigation rank the same models in the same order as the seeded story does.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories