Skip to content

Build1 publisher3 min readPublished

Anthropic widened its dual-use biology block on lost certainty about the old threshold

Anthropic's third misuse report says its 2025 Claude models sat well below the level that could help a sophisticated user with dangerous biology, and that for current models it can no longer make that assurance.

The Engineer · Build desk

Illustration accompanying Anthropic widened its dual-use biology block on lost certainty about the old threshold

What happened

  • Anthropic said Thursday it blocked attempts to use its models for cyberattacks, surveillance and research that could have led to biological weapons, in its third misuse report since March 2025.
  • Only one case in the report involved the newer Fable or Mythos-class models, a campaign Anthropic called an industrial-scale covert effort to extract a model's capabilities and replicate them elsewhere.
  • Publication came two days after an Anthropic researcher announced he was resigning over concerns that the company and its competitors are not acting responsibly in AI development.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint Dual-use biology prompts that the 2025 models answered now meet a wider restriction on Claude Fable 5, which puts vaccine and virology drafting work in the same blocked category as weapons research.
  • decision Picking a model version now also picks a refusal surface, so any team with life-science prompts has to re-test the same prompt set against each generation before it migrates.
  • contradiction The stated basis for the tighter filter is lost certainty rather than a measured capability jump, so customers are governed by what Anthropic says it cannot rule out about its own models.
  • precedent Anthropic wrote the voluntary format for abuse disclosure that grades its own products, and the format is now public, with code and prompt snippets included.

The threshold argument decides which prompts an integration now sees refused. Anthropic's report says its 2025 models, Claude Opus 4 and Claude Sonnet 4.5, "were well below the threshold where they could meaningfully assist a sophisticated user in carrying out dangerous biological research" [13]. The safeguards were built to that estimate. "As a result, safeguards on these models were less stringent, directed mostly at preventing access to content that might uplift novices in recreating known bioweapons," the report said [14]. For the current generation the company withdraws the claim: "But for today's models - which are capable of assisting in a range of complex scientific research tasks - the evidence is no longer certain, and we cannot make that same assurance" [15]. Newer models such as Claude Fable 5 carry "stronger safeguards that restrict access to a wide range of dual-use biological research queries" [16].

So the blocked category moved from recreating known bioweapons to dual-use research, and the published example shows how wide that is. The deliverable Anthropic refused was a funding proposal [9]. "The work discussed in the application involved gain-of-function research (that is, research that genetically alters an organism to create a new or enhanced biological property) on the chikungunya virus. This gain of function research was aimed at the virus' transmissibility and immune evasion properties," the report said [10]. Chikungunya is mosquito-borne and causes severe pain and fever [18]. Anthropic said such work could "certainly" be used to develop better vaccines and treatments, and that "it could also be used to make the pathogen more dangerous" [11].

The observation window runs from December 2025 to August 2026, nine months [8][19], and covers spyware vendors, "politically motivated individuals" and state-sponsored groups spreading propaganda [8]. Only one case touched the newer Fable or Mythos-class models, a case Anthropic described as "an industrial-scale, covert campaign to extract a model's capabilities and replicate them in another model without authorization" [12]. The company's own caveat limits what the sample is worth as a base rate. "The cases we share here aren't typical misuse, but rather examples of the most notable and novel threat activity we've identified to date," Anthropic said [3].

For anyone pricing the wider filter, the published account carries no count of blocked requests and no rate at which legitimate research trips the dual-use rules [21].

Anthropic frames the disclosure as its own obligation. "We're publishing this work because we believe we have a responsibility to disclose malicious misuse of our services. As models become increasingly capable, their risks will increase, unless AI developers and society's defenders act to make them safer," the company said [4]. The report prints snippets of the malicious code and prompts it found, and urges governments and AI competitors to identify and prevent similar abuse [5]. In the account, all of this is voluntary: no regulator requires it [22]. John Thickstun, an assistant professor of computer science at Cornell University, said it is an uncomfortable position for companies like Anthropic and OpenAI to be in when they are expected to determine what is safe versus unsafe behavior [17].

Anthropic also said elaborate cyberattacks no longer require sophisticated skills, and that lone individuals can now create threats that would not have been possible a year ago [20]. The report came out two days after one of its researchers announced he is resigning over concerns that Anthropic and its competitors are not acting responsibly in AI development [7]. The company is planning an initial public offering this fall [6].

What to watch

  • Whether Anthropic publishes refusal volumes or false-positive rates for the dual-use biology safeguards alongside the case studies.
  • Whether any regulator cites this report's format as expected practice for frontier developers. So far the account shows no such citation.
  • Whether the fourth report logs cases on Fable or Mythos-class models beyond the single distillation campaign.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories