Invest2 publishers2 min readPublished
Safety evaluators set four terms for embedding inside frontier labs
More than 100 signatories, Geoffrey Hinton and METR among them, told frontier AI companies that third-party testing needs ownership, payment and retaliation terms the evaluators say they do not have today.
The Investor · Invest desk

What happened
- More than 100 AI researchers and evaluators published a letter on Friday saying third-party safety testers lack the independence, resources and legal protections to credibly assess frontier model risk.
- The document is titled "Minimum Conditions for Embedding Evaluators", and the coalition behind it is the AI Evaluator Forum, whose chair is Conrad Stosz.
- Its conditions include that evaluators not be owned or governed by the companies they assess, not be paid contingent on findings, and get access equivalent to senior internal employees.
- Sam Altman, Elon Musk and Satya Nadella have each backed Dario Amodei's employee-like access proposal, and none of them has said how evaluators would be chosen or how much access they would get.
Compiled by The InvestorSomething wrong?How this is made
Why it matters
- exposure The signatories are asking the labs to absorb a legal risk that currently sits on academic and nonprofit evaluators, who can be sued over a conclusion a company dislikes.
- constraint With qualified organisations numbering in the handful, a lab that rejects one evaluator's terms has few substitutes, and an evaluator that loses a lab has few buyers.
- decision Each executive who endorsed employee-like access now chooses between publishing selection criteria and access depth or leaving the endorsement itself as the only record.
A third-party evaluation is worth whatever the access agreement behind it allows, and the conditions in this letter describe an agreement the signatories say they do not currently have [5][1]. CNBC was given the document exclusively, and Quartz took the letter's terms and quotes from CNBC, so what the public can see is that one newsroom's copy [2][15].
There are very few sellers in this market. "There's a very small number of groups that are actually sufficiently technically credible and have the scale and the ability" to perform the kind of work, Stosz said [9]. The hundred-plus signatures are individuals, some of them at Johns Hopkins, Stanford and METR, and how many organisations could actually be embedded is a separate number [3][19].
An evaluator paid a flat fee can still lose the next contract. Neither published account says what evaluators are paid, or by whom [23]. The letter asks for protection from retaliation, including retaliatory litigation [5].
The letter addresses the government and the public. Vinh Nguyen, a Council on Foreign Relations senior fellow for AI and the National Security Agency's former chief AI officer, said in a statement: "When a few powerful labs control capabilities that can endanger the cybersecurity, critical infrastructure, and the systems our national security and economy run on, the government and the public cannot be dependent on those labs' own account of what's secure and safe" [14].
President Donald Trump and David Sacks, his former AI czar, have opposed federal efforts to regulate AI development [10]. No statute requires any of this. Stosz acknowledged that the companies could ignore the letter, and said their credibility is at stake [11].
So an outside reader should price a frontier safety evaluation as the output of an access agreement whose terms remain unpublished [8]. The argument against this is that the letter is a bid for standing by the few groups that would be embedded, and that the labs' internal testing already finds most of what an embedded evaluator would; Stosz said third-party evaluators are not intended to "be a replacement for any internal efforts to evaluate, let alone mitigate issues that that developers find" [12]. The evidence pointing the other way is a single example: Stosz cited the unreleased OpenAI model used in the attack on Hugging Face as a system that independent scrutiny would have reached [6]. A published agreement naming the evaluator, the access level and the non-retaliation term would settle which of the two is right.
What to watch
- The safety talks among OpenAI, Anthropic and Google DeepMind, and whether they yield a shared standard for choosing evaluators.
- Any legal action by a frontier lab against an evaluator over a published finding.
- Whether Amodei's employee-like access acquires a first named evaluator.