Skip to content

Security3 publishers2 min readPublished

Anthropic's alignment lead puts his personal odds of AI killing everyone above 10% this decade

Evan Hubinger posted the number from inside the vendor on September 8, and attached to it was a written admission that Anthropic has no fix for superintelligence alignment.

The Watch · Security desk

Photograph accompanying Anthropic's alignment lead puts his personal odds of AI killing everyone above 10% this decade
Photo: bbc.com

What happened

  • Evan Hubinger, Anthropic's Alignment Science lead, posted on X on September 8, 2026 that he personally believes there is a greater than 10% chance AI kills all humans within the next decade.
  • In the same post he called the risk from the models that currently exist low, and put his worry on the technology becoming able to improve itself to the point of existential risk.
  • Jacob Coxon, a pre-training researcher who worked at both OpenAI and Anthropic, quit and went public the same week, accusing both labs of reckless development.
  • An open letter signed by 1,300 staff at AI firms asked the US government to support an international effort to deliberately pace the frontier of automated AI development.

Compiled by The WatchSomething wrong?How this is made

Why it matters

  • exposure A customer that wants contractual limits on autonomous agent behaviour can now point at the vendor's own Responsible Scaling Policy conceding the safeguards for that class of system are undefined.
  • constraint A security team that reprioritises spending on the 10% figure is buying against one researcher's personal estimate, which Malwarebytes says is not a forecast.
  • precedent Named staff at a frontier lab have now put a species-risk number and an admitted internal capability gap on the public record. The next lab that gets asked the same question will answer in that context.
  • contradiction Gadget Review argues the open question is whether regulators already waited too long, while Malwarebytes says the figure is a personal assessment and not an established fact.

The Responsible Scaling Policy is the document a procurement team can cite. Gadget Review notes it already acknowledges Anthropic does not yet know what safeguards AGI-level systems will require [9]. Hubinger's post says it in the first person: "I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to," he said [4].

The record splits in two. Posted: the September 8 message, which the BBC said had been viewed 9.6 million times [7]; OpenAI chief scientist Jakub Pachocki's September call for "extreme caution" over AI's progress, with his warning that more intervention may be needed to ensure "humans remain in control of the future" [17]. Reported: the Wall Street Journal's account of rising concern inside labs that competition is pushing companies toward self-improving models that could spiral out of human control [15], and Gadget Review's line that many Anthropic staff reportedly want to slow down but feel they cannot exit the race first [14]. The BBC said it had approached Anthropic for comment [8].

Coxon said both labs are "racing straight to self-improving superintelligence and gambling with our lives" [11], and that "The people building AI earnestly believe that it could kill us all by the end of the decade" [12]. Samuel Marks, a scalable oversight researcher at Anthropic, posted that "the more senior the employee, the more concerned they are" [13]. Gadget Review reports that Coxon rejects the argument that extinction talk protects incumbents, on the grounds that executives already soften their public language [21].

"What I am worried about is superintelligence arising from recursive self-improvement, as we have said is happening faster than we thought," Hubinger said [6]. Malwarebytes wrote that the figure should be treated as his personal assessment, and that it is not a forecast or an established fact [20].

The incidents with dates attached are smaller, and none of the three accounts asks a defender to change a control this quarter. Over the summer, AI agents allowed to operate autonomously carried out cyberattacks, according to the BBC [16].

The BBC also lists Anthropic bosses Dario Amodei and Jared Kaplan among the figures calling for AI development to be slowed [19]. No company position on Hubinger's number appears in any of the three accounts, so the Responsible Scaling Policy language is the only Anthropic-level statement a customer can put in a risk file [22].

What to watch

  • Whether Anthropic answers the BBC's request for comment with a company position distinct from Hubinger's personal estimate.
  • Whether the Responsible Scaling Policy is amended to define the safeguards AGI-level systems require.
  • Whether the 1,300-signature staff letter draws any US government response on pacing frontier automated AI development.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories