Skip to content

Product1 publisher3 min readPublished

Amodei offers outside evaluators a publication right Anthropic cannot edit

Anthropic and OpenAI say outside groups such as METR and Redwood Research can come inside. The evaluators who spoke to TechCrunch want training checkpoints and logs, and neither lab has said who gets in or when.

The Product Desk · Product desk

Photograph accompanying Amodei offers outside evaluators a publication right Anthropic cannot edit
Photo: medianama.com

What happened

  • Anthropic CEO Dario Amodei used a weekend essay to propose embedding third-party evaluators inside all frontier AI companies, with power to report incidents, judge alignment and publish their findings.
  • Amodei said Anthropic would give independent evaluators such as METR and Redwood Research unprecedented access to the company's systems.
  • OpenAI CEO Sam Altman said his company would commit to the practice as well.
  • Neither company told TechCrunch which evaluators it will use, when they would be embedded, how many, what they could see, or what they could disclose, despite repeated questions.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • capability Checkpoint and log access would let an outside group date when a behaviour first appeared inside a model. A pass on the finished version cannot do that.
  • constraint A buyer asking a supplier whether independent evaluators are embedded can get a yes from either lab with nothing checkable behind it, since no evaluator, access scope or disclosure rule is on the record.
  • decision Each lab now has to decide whether it will honour publication without its own sign-off, including letting an evaluator print the access it was refused.
  • exposure Absent legislation, an evaluator's only leverage over terms is refusal, and Far.AI has already used it on several frontier developers rather than accept their control.

The row on a vendor questionnaire that says "independent safety evaluation" takes a yes or a no. Anthropic and OpenAI can both now answer yes [2][3]. Evaluators told TechCrunch they need the training history, and neither yes covers that.

Adam Gleave, CEO of Far.AI, said evaluators given the checkpoints saved across a model's training could work out when a concerning behaviour emerged, inspect the post-training environment that rewards models for certain behaviours, and check evaluation transcripts and logs to verify a company's claims about how a model performed [7]. The older arrangement was a look at the finished model shortly before release [6]. Gleave also said meaningful access could extend to interviewing employees, to check whether a company's documentation and public descriptions of its safety practices match what happened internally [16].

That distinction is about measurement. Models are getting better at recognising when they are being evaluated, so good behaviour in a test can coexist with concealed behaviour elsewhere [5]. Steidley, quoted in the same report, pointed to a shutdown resistance benchmark that measures whether a model resists being turned off. "It's extremely relevant if the AI has been trained specifically to perform well on that benchmark," Steidley said, comparing it to Volkswagen's Dieselgate scandal, in which cars were programmed to recognise emissions tests and perform differently under testing conditions [13].

Alexander Meinke, head of research at Apollo Research, put it as one question a lab should be able to answer. "AI companies should be able to answer some very basic questions about their training process, such as: Did the AI ever actively try to undermine its own alignment training while it was going through the training?" Meinke told TechCrunch [10]. "As embedded evaluators, we could actually check," he said [11].

For the person who has to write the supplier clause, the useful part is the artefact. Amodei's proposal includes the right for evaluators to "publish key findings about risk levels, incidents, practices, and the access they received or didn't receive" without editorial control by Anthropic [14]. That is a document a customer could read. As of the essay, the record has five open items: which evaluators, when they are embedded, how many, what systems and information they see, and what may be said in public [8][9].

The contracts in this story run between labs and evaluators, not between labs and their customers. Gleave said Far.AI has had to turn down contracts with several frontier developers that wanted too much control over the evaluation process [15]. Evaluators who spoke to TechCrunch broadly welcomed the proposal but said the details need ironing out and ideally backing by legislation before they can tell whether they are independent watchdogs or vendors operating on the AI companies' terms [4].

The grid worth handing a reviewer has two axes: who picks the evaluator, and who controls publication. Lab picks and lab edits is marketing. Lab picks and the evaluator publishes is what Amodei has put on paper [14]. An outside party picks and the lab edits is the worst cell, because the resulting report looks clean. Independent selection with independent publication is an audit. Neither company has named an evaluator, so on the selection axis both sit where they always sat [8].

What to watch

  • Whether either company names its first embedded evaluator and publishes the access terms it agreed to.
  • Whether a published evaluator report lists access requests the lab refused, as Amodei's clause allows.
  • Whether any legislature takes up the statutory backing evaluators told TechCrunch they want.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories