Skip to content

Invest1 publisher3 min readPublished

A third-party platform automatically routed Anthropic's rejected prompts to another AI model

Anthropic banned accounts over five dual-use biology cases set out in a 154-page misuse report, said it is not alleging the researchers intended harm, and found that one refused request was routed on to another AI model.

The Investor · Invest desk

Photograph accompanying A third-party platform automatically routed Anthropic's rejected prompts to another AI model
Photo: foxbusiness.com

What happened

  • Anthropic said Thursday it disrupted several attempts this year by researchers to use its models for biological research that could aid weapons development, banning accounts and strengthening safeguards.
  • The 154-page report, "Detecting and Countering Misuse of AI", sets out five case studies of researchers using Claude for advanced biological research with potential dual-use applications.
  • Anthropic said sophisticated threat actors know providers are trying to detect dangerous uses and exploit the dual-use nature of biology to keep a kind of plausible deniability about their research.
  • A separate section describes an Iran-nexus actor that used Claude to analyse public data for targeting recommendations against US naval forces; Anthropic banned the account and shared threat intelligence.

Compiled by The InvestorSomething wrong?How this is made

Why it matters

  • exposure Legitimate pathogen work is now exposed to a ban that does not require proof of intent, and the resulting findings travel to government authorities and to other AI labs.
  • constraint Detection spend buys enforcement at the account boundary only: a broker sitting between the customer and the model can send a refused prompt to a different provider.
  • cost Investors get nothing to model from this disclosure, because the account attaches no spend, headcount, or lost revenue to the detections, bans and relay takedowns.
  • precedent By handing case detail to rival labs and to government, Anthropic sets the expectation that a frontier provider discloses its own misuse findings, and every provider that does not now stands out.

The routing matters more than the ban. Anthropic said it blocked the chikungunya grant-proposal request and then discovered the researchers were using a third-party platform to bypass regional restrictions and automatically route rejected prompts to another AI model [4]. The company's account, as reported, does not say what the second model did with them [4]. The refusal bound one supplier in a chain the customer had assembled.

The report runs to 154 pages and sets out five biology case studies [2]. Four of the subjects are described: chikungunya, bird flu, orthopoxvirus, and non-transmissible venoms and toxins [3][5]. One case's subject does not appear in the account [14]. Five worked examples show that a policy exists and that staff act on it, and they are not a measure of how much misuse the policy catches: the report gives no total request volume, no count of accounts banned, no rate at which the detections fire [15].

For an enterprise buyer, the operative term is the standard for a ban. Anthropic emphasised it is not alleging the researchers intended to cause harm [6], and the report states that "Biological capabilities are dual use: they can be used for beneficial or harmful purposes, and it is often difficult to distinguish between them" [7]. So the pattern of the requests is what closes the account, and Anthropic said it "banned all associated accounts, worked with partners to take down the relay networks that evaded regional blocks and shared our findings with affected AI labs and government authorities" [8]. A vaccine programme whose prompts are hard to distinguish from the harmful version carries two risks: losing the account, and appearing in material handed to competitors and to government.

The report arrives days after a senior Anthropic safety researcher said, according to Fox Business, that AI has a greater than 10% chance of "kill[ing] all humans" within the next decade [11], answering a former researcher at both OpenAI and Anthropic, Jacob Coxon, who wrote on X that "I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives" [12]. A company whose own staff put catastrophe in double digits has a commercial reason to publish 154 pages showing it can see the traffic.

Two findings would break the read that detection is now a fixed cost of frontier supply. If the provider that took the rerouted prompts publishes nothing and keeps the revenue, the sharing Anthropic describes is a one-way cost with no effect on the market it names [8]. If the next Anthropic report drops the case studies, the format was a one-off. Anthropic could not immediately be reached by Fox Business for comment [13].

What to watch

  • Whether the provider that received the rerouted prompts publishes its own account of them or keeps the traffic quietly.
  • Whether Anthropic's next misuse report gives a denominator: total flagged requests, accounts banned, and false positives.
  • Whether enterprise Claude agreements start naming the sharing of findings with other AI labs and government authorities.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories