Skip to content

Product1 publisher3 min readPublished

A grant draft for chikungunya gain-of-function work tripped Anthropic's biology classifier

Anthropic's misuse report walks through five biological case studies, the safeguards that surfaced each one and the account bans that followed. The scientists and labs involved stay unnamed.

The Product Desk · Product desk

Photograph accompanying A grant draft for chikungunya gain-of-function work tripped Anthropic's biology classifier
Photo: foxbusiness.com

What happened

  • Anthropic said in a report shared on Thursday that it had flagged and stopped multiple scientists who were using its AI models to create potential biological weapons.
  • The full report also documents the company's models being used as a surveillance tool and to create software exploits, propaganda and weapons systems.
  • Everyone found misusing Claude or Anthropic models had their account banned, and the company says the investigation findings are feeding into further safeguards.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • constraint With no banned-account count and no false-positive rate in the published account, a research organisation cannot size the odds that its own legitimate work gets cut off mid-project.
  • decision Because Anthropic counted a military research institute link against one request, labs now have to decide whose affiliation plus topic makes routine prompt text a shutdown risk.
  • exposure A scientist who loses access is left with nothing to answer publicly, since the report describes the conduct and withholds the name.
  • precedent Anthropic has given customers a concrete list of five harm categories and a named classifier to demand from every other provider they deploy.

Drafting a grant application is ordinary work for a working scientist, and the request that tripped the flag was a grant draft. Anthropic says its biological safety classifier fired in May on a request for Claude to author a grant for gain-of-function research on chikungunya virus [7]. The proposal described increasing the virus's transmissibility and its ability to evade immune response [9]. Anthropic says the fact that the proposed research was affiliated with a military research institute made it extra questionable [10]. The company adds that the virus has no licensed treatment and can cause debilitating symptoms for weeks or months [8].

Jacob Klein, Anthropic's head of threat intelligence, told The New York Times: "You are not seeing someone in a comic book kind of way say, 'Hey, I want to build a biological weapon to kill everybody,'" [5] "It's an incredibly nuanced situation," he said [6].

The company told the Times that detection is complicated because research into a new vaccine could look a lot like creating a bioweapon [3]. Anthropic says it erred on the side of being overly cautious in all cases because of the possible consequences [4]. That is a stated rule for breaking a tie, and the remedy attached to it is blunt: anyone misusing Claude or Anthropic models had their account banned [13].

Count the harm categories in the published account and there are five: biological misuse, surveillance, software exploits, propaganda and weapons systems [17][12]. What the account does not carry is a denominator. The report gives no count of banned accounts and no rate at which the classifiers fire on legitimate research [18].

Anthropic is also not naming the individuals or institutions its investigation connected to possible biological weapons work [14]. "The individuals implicated in these case studies are working scientists," the company says [15], and "We do not assert that they intended harm, and identifying them or their labs could expose them to harm" [16].

Two questions sort the workflows your researchers actually run. Does the text of the request resemble one of the biological case studies, including the bird flu gain-of-function work and the atlas of venom toxin peptides with "a generative pipeline that optimized toxin characteristics" [11]. And does the requester's affiliation add context a reviewer would read as adverse. The chikungunya case sits where both are true [7][10]. A peptide optimization pipeline at a commercial lab sits where the first is true and the second is not, and an overly cautious tie-break there costs a team its working account [4][13].

For that second case, the open questions for a provider are what fires and who inside your organisation hears about it. Anthropic's report answers the first for its own models, and the company says the findings are being used to improve safeguards and prevent future misuse [13].

What to watch

  • Whether a later Anthropic report publishes a banned-account count or a false-positive rate for the classifiers.
  • Whether any lab or scientist comes forward to contest a ban, given that Anthropic withholds identities.
  • Whether other model providers publish comparable case studies naming the safeguard that fired.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories