Skip to content

Product1 publisher3 min readPublished

Anthropic co-founder says a third-party-checkable kill switch may have to be mandatory

Jack Clark told the BBC that most labs including Anthropic can already pull the plug on a model, and he asked whether rules should require one and let a third party verify it. It is buyers who have to work out how much of the service that switch turns off.

The Product Desk · Product desk

Photograph accompanying Anthropic co-founder says a third-party-checkable kill switch may have to be mandatory
Photo: yahoo.com

What happened

  • US lawmakers have put forward legislation dubbed the Kill Switch Act, which would require companies to have a way to shut down problematic AI tools.
  • The UK government has rejected creating a kill switch, with a spokesperson saying it would not prevent AI systems being developed or misused elsewhere.
  • Trump has rejected attempts to slow AI down, posting on social media that AI taking over the world and destroying humanity is a hoax.
  • Clark spoke days after a viral post by an AI researcher who quit Anthropic over concerns that AI could wipe out humanity.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • exposure Under the bill as described, an agency order to limit or switch off a tool would reach the model provider, and every company reselling that model inherits the downtime without a seat at the table.
  • constraint There is no common interface to certify: labs each built their own way to pull the plug, so a verifier would have to qualify each provider separately.
  • decision Procurement gets a question it can ask today, before any statute exists: what is the smallest unit the vendor's switch can turn off, and who has tested that.
  • contradiction Anthropic is arguing for a slower, more monitored industry while reportedly readying a record-setting listing, and OpenAI's Altman has held his own float back over the same safety debate.

The question for whoever is shipping on a hosted model is narrow: what does "pull the plug" actually turn off. Clark said most labs, Anthropic included, have some way of doing it, and that policymakers may need to enforce having one [2]. His own framing was two questions. "Should you mandate for companies to definitely have a killswitch? Is that kill switch verifiable by a third party?" he asked [3]. "I think that's the kind of thing society is going to want to know and might want to eventually pass rules around," he said [4]. Clark told the BBC that the specifics of requirements and verification should be part of "the larger policy conversation" [5].

Verification is the harder half. Each lab built its own way to pull the plug [2], so an outside auditor would have to learn each one, and a team running two model providers for failover cannot assume the two switches cut at the same layer. A vendor can prove it is able to stop serving a model worldwide and leave its customer guessing whether the shutdown is scoped to the model or to one tenant's traffic.

The bill in Congress would also give certain government agencies the power to demand that a tool be turned off or limited [7]. The order goes to the model provider. Only the provider is party to it, and the outage lands on its customers.

The extinction percentages are the part of this that travels furthest and helps least with planning. Anthropic scientist Evan Hubinger said he personally thought the possibility of human extinction from AI was ">10% within the next decade" [13]. Geoffrey Hinton told the BBC on Friday that a 10% chance of AI killing all humans was "not unreasonable" [14]. Clark, asked for his own figure, said "I don't think these statistics are that useful", and said letting AI continue as a totally unregulated industry was a bad idea [15]. "We are rolling dice with immense risks," Clark said. "And the point is, we have to change the course of this industry" [16]. Of the three, two gave a number and one declined [21].

Dario Amodei called over the weekend for the pace of AI development to slow and be more closely monitored, and said any action to rein it in should come "without sacrificing commercial advantage" [10]. He did not say how development would be slowed [11]. Anthropic is reportedly preparing a potentially record-setting stock market debut [12]. Sam Altman said on Friday that OpenAI, valued most recently at $852bn, would not do the same this year because of the current debate around AI safety [19].

For a renewal conversation, two axes are enough. The first is scope: does the shutdown reach the whole service, or your deployment. The second is evidence: has anyone outside the vendor tested it. Scoped and externally tested is a control you can write a runbook against. Global and untested is a single point of failure that the vendor asserts works and nobody outside has checked, and the two mixed quadrants tell you which of the two questions to press first. Both Anthropic and OpenAI have self-reported incidents where AI agents acted in ways that were unexpected [17]. A mandated switch is for events like those. On the customer's side it arrives as an outage.

What to watch

  • Whether the Kill Switch Act's text names who performs verification, and at what scope the shutdown has to operate.
  • Whether Anthropic's reported IPO paperwork describes shutdown capability as a risk factor or as a control it already has.
  • Whether the UK government's rejection holds once a US mandate has statutory language attached to it.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories