Skip to content

Product1 publisher3 min readPublished

Anthropic sorts eight months of Claude misuse into seven named harm areas

Anthropic's threat intelligence team published case counts for biological misuse and state-linked surveillance of Claude, and did not say how long most of the logged activity ran before the company found it.

The Product Desk · Product desk

Illustration accompanying Anthropic sorts eight months of Claude misuse into seven named harm areas

What happened

  • Anthropic published a survey of the last eight months of attempts by threat actors to use Claude for what it calls malicious activity, covering use cases the company labels atypical.
  • The company sorted the activity into seven harm areas: cyber operations, influence operations, surveillance, scams and fraud, biological misuse, conventional weapons development, and distillation.
  • Anthropic alleges researchers, some of them state-supported, prompted Claude to write grant applications for experiments on the chikungunya virus and on orthopoxvirus, the family that includes smallpox.
  • One operation with suspected ties to the People's Republic of China used Claude and bulk data from monitored WhatsApp and Telegram chats to track, profile and recruit politically connected Uyghurs and journalists.
  • Anthropic attributes to a Russia-based freelance agent an account that used Claude Code to engineer a full-stack autonomous first-person-view kamikaze drone swarm.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • exposure The people profiled in the surveillance cases never held a Claude account. Banning the operator is the only enforcement the vendor has, and it does not undo the harvesting for the person whose messages were read.
  • constraint A case count with no accounts-reviewed figure and no elapsed time cannot carry a trend claim, so anyone comparing this report to next year's will be comparing detection effort and not misuse volume.
  • decision A trust-and-safety team can now prohibit named behaviours, grant-application drafting and control circumvention among them, instead of a general clause about harmful use. Naming them also commits the team to detecting them.
  • precedent Procurement has one published format it can ask other model vendors to match. The report says nothing about what any other vendor discloses, so there is no baseline yet to score the answers against.

At least 16 accounts tied to Iranian paramilitary and domestic security agencies used malicious browser extensions built by Claude to harvest data from 6,388 Iranians, according to Anthropic [15]. Divide the second number by the first and you get about 399 people per account, and because 16 is a floor, the real average is lower [17]. In the same section the company says an Iran-nexus threat actor used Claude to pinpoint U.S. naval targets [16].

In the biological cases the failure has two stages. Anthropic said the users "circumvented controls" designed to prevent individuals in countries without access to Claude from using the product, and also attempted to hide the purpose of their research to evade safeguards [6]. The company described the individuals in each case study as working scientists and did not name them [7]. "We do not assert that they intended harm, and identifying them or their labs could expose them to harm," the company wrote in its report [8]. It banned the accounts upon discovering the activity [9].

Mashable's account of the report counts five biological case studies [4] and nine surveillance instances spanning China, Iran, west African countries and for-hire networks [13], which is 14 enumerated cases in two of the categories [18]. Anthropic did not include details on how long the activity had been occurring in most cases [3]. Those are numerators with no denominator: no count of accounts reviewed, no time between first prompt and ban.

Anthropic argues the record itself is the new thing. "Historically, this kind of work has been uncovered by governments, United Nations panels, and outside investigators, who piece it together from recovered hardware and public sources," the company explained [10]. The China-based account in the weapons section fits that description of a gap: Anthropic says the user, possibly a military-industrial researcher, built electronic warfare modules with Claude work tools to defeat enemy radar and communications and suppress air defenses, in a simulation that included 12 targets in Taiwan [12]. That came out of account activity on a vendor's servers, and the vendor chose how much of it to print.

The next document of this kind, from any vendor, can be sorted on two axes. Does the item carry a denominator, and was the misuse caught during the session or after the fact. Items with both are a control test. Items with neither are a case list, which is good material for drafting policy and no use for measuring whether the policy works. Almost everything here sits in the second group, which is still more than a vendor safety page gives a buyer. The two figures that would move items into the first group are accounts reviewed and time from first prompt to ban.

What to watch

  • Whether the next edition publishes accounts reviewed and the time from first prompt to account ban for each case.
  • Whether any other model vendor publishes case counts in comparable categories; that would give a buyer a baseline to score them against.
  • Whether the geographic control that these accounts circumvented is changed, and whether Anthropic says so publicly.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories