Leadership1 publisher3 min readPublished
Anthropic's eight-month misuse report sorts disrupted Claude abuse into seven harm areas
Anthropic says it disrupted state-sponsored groups, spyware vendors and criminal fraud operations using Claude between December 2025 and August 2026. The published cases are the ones it calls most notable and novel.
The Board Room · Leadership desk

What happened
- Anthropic's September 2026 threat intelligence report covers misuse it disrupted between December 2025 and August 2026 across seven harm areas, including cyber operations, influence operations, surveillance, scams and fraud, and distillation.
- The actors it names include suspected state-sponsored groups, financially motivated criminals, commercial spyware vendors, state propaganda institutions and politically motivated individuals.
- Claude Haiku, Sonnet and Opus appear in the misuse cases, and Anthropic says no case involved its Fable or Mythos-class models apart from one illicit distillation case.
- Cases include a network of fake dating apps designed to defraud users and surveillance systems built to identify and monitor dissidents.
- Anthropic says it disrupted each operation, strengthened its safeguards on what it learned, and shared intelligence with authorities and industry partners where appropriate.
Compiled by The Board RoomSomething wrong?How this is made
Why it matters
- capability An acceptable-use policy can now be written against seven named harm areas instead of a generic misuse clause, and each one becomes a question a vendor can be asked to answer in writing.
- constraint Because Anthropic selected the cases for novelty, the report cannot support a frequency estimate or a comparison between vendors, so a risk register citing it as prevalence evidence is citing a curated set.
- exposure Organizations whose threat model assumed a well-resourced adversary are, on Anthropic's account, now reachable by individual operators holding comparable tooling.
- decision Model-class selection becomes a control decision, and the buyer has to make it without knowing whether the clean Fable and Mythos record reflects safeguards or lower adoption.
Anthropic has now published four of these reports: March 2025, August 2025, November 2025 and this one in September 2026 [10][1]. A security team treating the series as a monitoring feed should note the spacing, which runs five months, then three, then about ten [21].
The company is precise about how it chose what to publish. "The cases we share here aren't typical misuse, but rather examples of the most notable and novel threat activity we've identified to date," Anthropic wrote [9]. That sentence sets the document's limit as evidence. It supports a list of the shapes abuse takes; it cannot carry a rate. A skeptic would say a vendor publishing its own disruption record has an interest in the record looking managed, and the report concedes that "sophisticated and persistent threat actors continuously test our safeguards and try to circumvent the technical measures we use to detect and prevent misuse" [11]. What survives that objection is the part a buyer can convert into questions: seven named harm areas and five named actor types [2][22].
The cyber trend changes which control matters. Anthropic wrote that AI "has collapsed the labor and tooling gap that used to separate well-resourced, state-sponsored operations from individual operators" [12], and argued that "many commentators focus on the risk of AI developing exploits at scale. While this is a danger, the risk from AI adoption is more pronounced across the cyber kill chain, where adversaries can operate faster, across a broader and deeper surface area, with fewer resources" [13]. It measures that boost as "uplift", through speed, scale and depth [14]. One reported case involves a hacktivist using stolen API keys, leaving key custody and volume anomaly detection to the customer's own systems, not the vendor's [15].
Anthropic attributes the clean record for its Mythos class to safeguards that "greatly reduce its ability to perform harmful cyber tasks" [16]. Without usage volumes by model class, a reader cannot separate a safeguard that worked from a model fewer actors reached for.
There is also a dating problem for anyone building a timeline from this. The cyber section opens by saying "over the past six months, our Threat Intelligence team identified and disrupted a series of cyber operations in which threat actors used Claude" [17], then states that its cases span December 2025 through August 2026 [18], a period the report elsewhere calls eight months [3]. That is two months longer than the six-month window it opens with [23].
The decision available this quarter is narrow and concrete: whether an acceptable-use rule and a detection rule exist for each of the seven areas, distillation included, and which of the seven the vendor will say in writing that it monitors [2]. The consequence arrives with the next report, whenever it comes. Anthropic said it hopes the findings help other developers "recognize similar patterns on their own platforms," a comparison that needs logs a company was already keeping before it read the report [19].
What to watch
- Whether the next report arrives within a quarter or after another ten-month gap, and whether it carries incident counts beside the case studies.
- Whether any other frontier model vendor publishes a harm-area breakdown a buyer can set beside this one.
- Whether the single distillation case involving a Fable or Mythos-class model produces named parties or legal action.