Skip to content

Product1 publisher3 min readPublished

Anthropic and OpenAI pledge to embed outside safety evaluators with employee-like access

Dario Amodei proposed that frontier labs host independent evaluators inside the company, and Sam Altman said OpenAI will do the same. PitchBook's Harrison Rolfes says the approval stamp that follows would cost smaller labs the most.

The Product Desk · Product desk

Photograph accompanying Anthropic and OpenAI pledge to embed outside safety evaluators with employee-like access
Photo: yahoo.com

What happened

  • Anthropic CEO Dario Amodei proposed in an essay this month that frontier AI companies embed independent safety evaluators with employee-like access.
  • OpenAI's Sam Altman tweeted that committing to independent evaluators with employee-like access is a great idea and that OpenAI will do the same.
  • PitchBook senior analyst Harrison Rolfes says the evaluation costs could form a moat around the largest labs, because smaller labs and new entrants would struggle to shoulder the fees.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • cost On Rolfes's reading the fee attaches to whoever needs a release cleared, so an open-weights publisher pays a similar evaluation bill out of a far smaller budget than OpenAI's.
  • decision Build-versus-buy reviews on open models now need a row for who clears a release and what that clearance costs, and no party quoted has published a price to put in it.
  • precedent If the approval slate is agreed among labs rather than legislated, the people setting the bar can move it without a rulemaking, a comment period, or a vote.
  • contradiction The company that named the capture risk in 2024 is the one now proposing resident evaluators, which lets a smaller lab quote Anthropic back at Anthropic in a policy fight.

Amodei's proposal, as Fast Company describes it, is about access: independent evaluators sitting inside a frontier lab with employee-like reach, so the company is not the only judge of its own safety [2]. Anthropic's policy page says AI companies "shouldn't be the only ones deciding whether their systems are safe" [3]. PitchBook senior analyst Harrison Rolfes describes the result as permission. "It actually adds a very large cost because now you need a third-party stamp of approval to release models," Rolfes said [5].

Fees are one part of it. Evaluators also need computing resources, and the costs grow with the complexity of the model under test [13]. Marius Hobbhahn, CEO of Apollo Research, wrote in 2024 that "Since models are now capable enough of acting as LM agents, evals have to be increasingly complex tasks, which significantly increases the overhead per eval" [12]. Anthropic said much the same in a 2024 blog post, that the cost of evaluating huge and complex AI models "is high, and getting higher" [10]. No one quoted puts a dollar figure on a single frontier evaluation [22].

The body doing the regulating might not be a government. Fast Company reports the labs could instead agree among themselves on a set of independent evaluation firms, naming METR, Redwood Research and Apollo Research as candidates [8]. Rolfes compares that to the Big Four accounting firms, paid to vet other companies' financial statements [9]. A slate of approved evaluators agreed among the largest labs would be a standard the labs set themselves. Critics, in the same account, see the endpoint as an AI industry supplied by OpenAI, Anthropic, SpaceXAI, Google, Microsoft and perhaps Mistral in Europe [14], which is six companies, one of them outside the United States [20].

Wired's Maxwell Zeff reported on September 10 that OpenAI had asked members of Congress whether orchestrating an industry-wide slowdown in frontier AI development would violate antitrust law, citing people close to the company; that reporting does not say OpenAI asked about an agreement among industry players [19]. Two days later, Amodei's September 12 essay called for "pacing" frontier development so the labs could catch up on safety and alignment research [15][21]. Altman publicly agreed, saying "we need to pace the frontier" [16]. Musk's response was "Dario is right" [17]. OpenAI policy chief Chris Lehane later told reporters the company had already been working with Anthropic and Google DeepMind on AI safety for several weeks [18].

For a team deciding what to build on, whether you consume a model or publish one decides it. Call a hosted API and an evaluation regime arrives inside the vendor's price, invisible on your side. Ship weights that other people download and the stamp Rolfes describes lands on your release date, at whatever the evaluators charge [5]. For most roadmaps, what settles it is whether shipping requires an outside sign-off, and whether the schedule can absorb the wait for one.

What exists today is a proposal from one CEO and a public agreement from another [2][4]. No statute in this record requires a third-party evaluation, and the cost case rests on Rolfes plus Anthropic's own writing. Anthropic put that warning on the page in 2024: tests should be a "well scoped, small set", it said, "or else we'll create regulatory burdens and increase the possibility of regulatory capture" [11].

What to watch

  • Whether METR, Redwood Research or Apollo Research publishes a fee schedule or sample scope of work for a frontier model evaluation.
  • Whether Congress answers OpenAI's antitrust question about an industry-wide slowdown, and in what form.
  • Whether any open-weights developer commits to hosting embedded evaluators or publicly declines to.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories