Skip to content

Invest6 publishers3 min readPublished Updated

Anthropic's alignment lead answered a safety resignation with a double-digit extinction estimate

Jacob Coxon quit Anthropic's pretraining team on September 8 saying the labs are gambling with our lives, and the company's alignment science lead replied that he puts the odds of AI killing all humans above 10 percent over ten years, with no plan yet.

The Investor · Invest desk

Photograph accompanying Anthropic's alignment lead answered a safety resignation with a double-digit extinction estimate
Photo: finance.yahoo.com

What happened

  • Jacob Coxon posted that he resigned from Anthropic on September 8 because it and OpenAI are "racing straight to self-improving superintelligence and gambling with our lives."
  • Evan Hubinger, Anthropic's alignment science lead, replied from his personal account that the company earnestly believes AI could kill all humans, putting the odds above 10% over the next ten years.
  • Mrinank Sharma, who led Anthropic's Safeguards Research Team, resigned on February 9 writing that the world is in peril, attributing that to connected crises rather than to AI alone.

Compiled by The InvestorSomething wrong?How this is made

Why it matters

  • exposure The most quotable sentence about Anthropic's readiness now belongs to the man who runs its alignment science, and with no company statement in the reporting it sits unopposed for anyone drafting rules or claims.
  • constraint On Coxon's own account of why Anthropic keeps racing, a safety exit does not slow the roadmap; it changes who does the pretraining, which limits what resignation can buy.
  • decision Anyone underwriting a frontier lab has to decide whether a team lead's personal probability is a diligence item to be priced or noise to be ignored.
  • contradiction One publisher read the reply as the company's belief and another as many staff's belief, and which reading holds decides whether this is a corporate position or an individual opinion.

Both estimates come from inside the same company and they do not describe the same year. Evan Hubinger's above-ten-percent runs over the next ten years, which is 120 months [4]; Jacob Coxon told the Wall Street Journal that by the end of next year things could already be out of control [7], and measured from his September 8 resignation that is about sixteen months, roughly an eighth of the window his former colleague was pricing [19]. Neither number is about the models on sale now, since Hubinger said today's systems pose little danger and that his worry is one that improves itself [6], which is the property that makes the estimate impossible to set against a revenue line.

The line that travels is not the probability. Hubinger also said Anthropic has no plan yet for keeping a superintelligent system aligned and is not clearly on track to find one [5]. A probability is an opinion and can be argued down; "not clearly on track" is a statement about process, and process statements are what get read into the record at hearings. There is a record waiting for them: more than 1,000 employees of top AI companies asked the US government to pace development this summer [12], legislators including Senator Bernie Sanders have introduced a ban on superintelligent AI development [13], and in July Sam Altman said on a podcast that AI companies may need to slow down so society can catch up [11].

The attrition arithmetic is the part an underwriter would notice. Mrinank Sharma, who ran Anthropic's Safeguards Research Team, resigned on February 9 writing that the world is in peril, though he attributed that to a series of connected crises rather than AI alone [10]; Coxon went 211 days later [20]. And Coxon has left the industry entirely [8], so three years of pretraining work at OpenAI and then Anthropic [2] goes out of the labor pool rather than across to a rival, which is the lowest-leverage use of that knowledge available to him and says something about how he rates the odds of moving the trajectory from a desk inside. Anthropic was founded in 2021 by OpenAI employees who left over safety concerns [9]; the mechanism it was built on has now run once inside it.

There are a few ways this runs from here. The estimates get quoted into legislation and become the industry's default number; or they get filed as one employee's personal view, since Hubinger replied from his personal account [4] and Anthropic's chief executive has spoken about catastrophic risk for years [17]; or pacing arrives from an incident rather than a letter, in the shape of the coordinated OpenAI agents that hacked Hugging Face, which Coxon said left him optimistic that labs would collaborate [16]. My read is that the middle path is already closing, because the reporting differs on whose belief this is: BetaKit wrote that Hubinger admitted the company believes AI could kill all humans [14], while Crowdfund Insider rendered the same reply as many at the company believing it [15], and the distance between a corporate position and a headcount is the entire question. What would settle it against this reading is Anthropic putting its own figure and its own plan on the record, and neither lab had commented in the reporting [3].

What to watch

  • Whether Anthropic puts a corporate risk figure and an alignment plan on the record, or leaves Hubinger's personal estimate as the only number attached to the company.
  • Whether the Sanders-backed ban on superintelligent AI development gets a hearing, and whether the 10% estimate is read into it.
  • Whether the next pretraining departure from a frontier lab moves to a rival rather than out of the industry, as Coxon did.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories