Skip to content

Leadership1 publisher3 min readPublished

Amodei gives his botnet warning a 6-12 month deadline

A sharptext.net column argues the loudest AI risk estimates should be scored against the record of the people issuing them. Of the four forecasts it assembles, only one can be checked inside a year.

The Board Room · Leadership desk

Illustration accompanying Amodei gives his botnet warning a 6-12 month deadline

What happened

  • A sharptext.net column reports that an Alignment Science lead at Anthropic puts the chance his company's technology leads to human extinction in the next decade at greater than 10%.
  • It sets that estimate beside Anthropic chief executive Dario Amodei's own stated 10-25% chance that AI will have a catastrophic impact.
  • The column recalls that in 2019 Amodei and other OpenAI researchers believed GPT-2 was too dangerous to release to the world.
  • On the Hugging Face incident, it says the behaviors were not emergent, the reporting from safety organisations was arguably too dramatic, and the fallout was exceptionally limited.
  • Its author argues the frontier AI community should be treated as fringe activists who are sometimes right but often wrong.

Compiled by The Board RoomSomething wrong?How this is made

Why it matters

  • decision A forecast with a 6-12 month window can be entered in next year's review and marked right or wrong. A probability spread over a decade never comes due inside a planning horizon. The two cannot be treated as the same input to a budget.
  • constraint Both probability figures a board is likely to be handed come from one company's payroll, so citing them together does not give it two independent estimates of the same risk.
  • contradiction The column doubts the messengers while conceding the capability: it expects recursively self-improving AI soon and disputes only volition and recourse. Those doubts about the messengers do not support deferring technical controls.
  • cost Managers absorb the work of answering staff who have heard the extinction figure without its time frame, and what this column gives them to answer with is an argument about credibility.

Anthropic's chief executive put a clock on his own warning. Writing last Saturday, Dario Amodei said he worries that in 6-12 months an agent swarm could be capable of taking over the entire internet with a persistent botnet, potentially causing hundreds of billions of dollars in damage, and that the scale of damage would keep increasing from there if AI becomes more powerful without the necessary guardrails [1].

Only the botnet forecast can be scored without waiting a decade. Of the four Amodei forecasts the column assembles, three carry dates: 2019 on GPT-2, last year on employment, last Saturday on the botnet, alongside an undated 10-25% chance of catastrophic impact [18]. The extinction estimate the column attributes to an Anthropic Alignment Science lead runs out to the next decade [2]. A board approving security spend this quarter can put the botnet forecast back on the agenda next autumn. The other three sit outside any cycle a board runs.

The column's argument about whom to trust has two halves, and they are not equally usable. One half is record and incentive: the author writes that today's AI leaders "have their own spectrum of social and financial incentives, and most importantly, they are not scientific authorities worthy of deference by default" [7]. The other half is character. "some people are just crazy. And crazy people are not worth indulging if you'd like to stay sane," he wrote [8]. Anyone can check a track record against the dates. Judging temperament is another matter, and an employee who has just read a double-digit extinction probability gets no answer from being told that the person who said it is unbalanced.

On specific scenarios the column argues physical feasibility. Building a bioweapon is far harder in practice than the scenario assumes [14]. Nuclear launch chains use multi-factor authentication before launch, almost all of it involving multiple humans, and even Russia's dead hand program reportedly has to be activated by a person [12]. Most large companies and financial institutions use air-gapped backups against exactly the records-destruction case, the author writes, adding that not all of that work is public [13]. Two of those three are answerable inside a single firm: whether the backups are air-gapped, and who can reach them.

On capability the column agrees. Its author writes that he is sure some version of recursively self-improving AI will exist soon if it does not already, and puts his disagreement on whether such agents would have volition or malicious designs of their own, and whether humans would have recourse if they did [11].

Both probability figures come from one payroll [19]. An Anthropic alignment lead supplies the extinction number and Anthropic's chief executive supplies the 10-25% range, so a reader who cites the pair has cited one employer twice. The author describes himself as a relative tech outsider who nevertheless pays close attention to this world [15]. A Pod Save America host wrote this week that he takes very seriously the warnings of the people who are actually building the technology, and the column says he has it precisely backwards [9].

What to watch

  • Whether the persistent-botnet forecast holds by the end of the 6-12 month window Amodei set last Saturday.
  • Whether the Anthropic alignment lead's greater-than-10% extinction estimate is ever published with a stated method behind it.
  • Whether safety organisations' reporting on the next containment incident matches the limited fallout the column describes at Hugging Face.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories