Skip to content

Product1 publisher3 min readPublished

Anthropic's bioweapon filters skipped 133 million contractor chats for eleven months

Its own Risk Report says an internal flag that disabled blocking also disabled logging, on a surface staffed by vendors that could not screen out CB-1 threat actors.

The Product Desk · Product desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Photograph accompanying Anthropic's bioweapon filters skipped 133 million contractor chats for eleven months
Photo: thenextweb.com

What happened

  • Anthropic published its Risk Report on 14 August, covering the period to 15 July.
  • Anthropic runs blocking classifiers that are meant to stop a model helping anyone build a biological weapon, and first deployed models carrying those safeguards in May 2025.
  • From May 2025 until April 2026, the blocking classifiers did not run on any traffic through Anthropic's human feedback platforms; the report's section describes this as eleven months with the filters off.
  • Roughly 50,000 people had access to the affected human feedback platforms.
  • Those people generated around 133 million exchanges during the affected period.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

Anthropic's Risk Report, published on 14 August and covering the period to 15 July, discloses that the blocking classifiers built to stop its models helping anyone construct a biological weapon did not run on any traffic through its human feedback platforms from May 2025, when models carrying those safeguards were first deployed, until April 2026 [1][2][3]. Roughly 50,000 people had that access and generated around 133 million exchanges, and the vetting of those people was done by outside vendors rather than by Anthropic [4][5][6]. That is about 2,660 conversations per contractor on a surface with the safety layer absent [1].

The mechanism is the transferable part. According to the report, a flag intended only for internal use switched off the blocking behaviour and switched off the logging with it, and flagged traffic "was not recorded or propagated to any review mechanisms" [7][8]. Enforcement and observability shared one control, so the gap could not be found afterwards without going back to raw transcripts [8].

The vendor layer is the other part. Anthropic writes that many of the vendors "did not have screening processes capable of stopping even CB-1 threat actors", and that the vast majority of contractors could hold open-ended conversations rather than simply rate a fixed set of answers [9][10]. A footnote states that before April 2026 a threat actor could probably have been hired at one of its vendors, and that a red-teaming role would have done it [11].

The same platforms produced a second incident. An outside tip in April 2026 led Anthropic to confirm that a few contractors at data-labelling vendors had exploited a flaw to obtain an API key and used models outside their assigned work, including Mythos Preview, among its most capable [12][13]. The access path stayed open for several weeks, with Mythos Preview inside it for roughly two of them and running without blocking biological classifiers [14]. Anthropic contained it within 90 minutes of learning of it and closed the vector the same day, and says no model weights were taken, no customer data was reached, and its core networks were not breached [15][16].

The retrospective review is thinner than the exposure. Anthropic ran Claude Sonnet 5 over every human turn from the affected period to flag harmful biological content, which surfaced 1,197 high-risk transcripts, of which 757 came from Anthropic's own teams on the same infrastructure and all but 62 of the remainder came from deliberate red-teaming [17][18][19]. That leaves 378 red-teaming transcripts and 62 others [2]. Staff read 92 transcripts in total, the 62 plus 30 randomly chosen red-teaming ones, and found no clearly concerning misuse but a handful of potentially dual-use conversations [20][3]. Anthropic says it is very unlikely the gap raised real-world risk, partly because the conversations were mostly short [21].

Then the sentence operators should copy into their own risk registers: the discovery "leads us to believe that there is an increased likelihood of other, similar issues unknown to us" [22]. Anthropic's February report was published with the gap open and, it now writes, "did not consider our human feedback platforms as a risk surface"; the company has revised its February assessment from very low to low [23][24]. Most coverage instead led on the misalignment rating, which Anthropic also moved from very low to low while noting its own arguments "likely still support a designation of 'very low'" and attributing the change to uncertainty rather than new evidence [25][26][27].

What to watch is whether any vendor of frontier models will supply a per-surface coverage map for its safeguards, rather than a model-level claim. The practical questions for a buyer are which internal flags can disable enforcement and logging together, who screens humans holding open-ended model access, and how a gap would be detected at all once the logs are gone [8][9]. Also worth watching: whether the next report re-rates anything else retroactively [24].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories