Skip to content

InvestNot yet confirmed elsewhere1 publisher3 min readPublished

OpenAI and Anthropic turn to small nonprofit evaluators whose funding is still unsettled

Anthropic and OpenAI are turning to a handful of small, mostly nonprofit evaluators such as METR, Apollo Research and Transluce to check their models. With no federal push for rules, who pays them, what they can see and whom they report to are still open.

The Investor · Invest desk

How we use AISend a correction

Illustration accompanying OpenAI and Anthropic turn to small nonprofit evaluators whose funding is still unsettled
Generated illustration
Who funds AI evaluators, and what they see, is unanswered Where each party in CNBC's reporting stands on bringing outside evaluators into AI labs.

Evaluators: funding, access and reporting are unanswered. Anthropic: CEO pledged to embed evaluators. OpenAI: finalizing assessor contracts. Balesni and Korbak: OpenAI cites policy breaches; they cite evaluator contact. Trump: voluntary accord urges using an evaluator.

Who funds AI evaluators, and what they see, is unanswered
WhoHowKindClaim
Third-party evaluatorsHow they are funded, what access they get and what the reporting structure looks like remain unansweredconstraint12
AnthropicCEO Dario Amodei pledged last month to embed independent evaluators in the companydecision6
OpenAISays it is finalizing contracts with third-party safety assessors and will announce details in the coming weeksdecision8
Balesni and KorbakOpenAI says they broke sensitive-information policies; they believe it was over how they communicated with evaluatorscontradiction15
President TrumpVoluntary accord from late September encourages AI companies to partner with an independent auditor or evaluatorprecedent7

What happened

  • Anthropic CEO Dario Amodei pledged last month to embed independent evaluators inside his company, and OpenAI CEO Sam Altman quickly endorsed the move.
  • OpenAI fired three employees last week, citing its rules on sensitive information, and two of them say the real reason was how they talked to third-party evaluators.
  • OpenAI disputes that account and says it is finalizing contracts with third-party safety assessors, with details due in the coming weeks.

Why it matters

  • exposure Until a party other than the labs pays, the only clients on record are the companies whose models the evaluators grade. Their independence depends on contract terms that have not been published.
  • contradiction OpenAI calls close work with evaluators core to its safety effort, yet two fired staff say contact with evaluators cost them their jobs. How that dispute ends decides whether lab employees will talk to outside checkers at all.
  • precedent With a voluntary accord backed by the president and no federal push for rules, the access and reporting terms OpenAI and Anthropic set now become the template other labs start from.

"To a degree, the problem, as always, is money," Suresh Venkatasubramanian, a computer science professor at Brown University, told CNBC. "Who is paying for these companies to do their work? How are they going to support them? You need an ecosystem, you need a viable business model for this." [5]

Most of the groups he is talking about are nonprofits [2]. So far their main job has been to assess what models can do and what risks they carry, and to flag cases where a model behaves badly [2]. CNBC reported that the groups are still getting established, even as capital pours into the industry at historic rates and new models are released at an unprecedented pace [3]. The report does not give a budget or headcount for METR, Apollo Research or Transluce.

So far the only named buyers are the labs being checked. OpenAI said on Friday it is "actively finalizing contracts with third-party safety assessors and will announce details in the coming weeks" [8]. A spokesperson said that work builds on existing collaboration with METR and Redwood Research [9]. Anthropic's chief executive started this round with his pledge to embed evaluators, and the company did not respond to CNBC [6][10]. The White House is offering encouragement. President Donald Trump praised AI executives for their "tremendous self-policing" and, in a voluntary accord presented in late September, urged them to "partner with an independent external auditor or evaluator" [7]. For now, CNBC reported, Anthropic, OpenAI and the infrastructure partners profiting from the boom are writing the rules [13].

The money has three obvious places to settle. Lab contracts could give a small evaluator defined access and a reporting line outside the company, so lab money pays for a check the lab cannot edit. The work could move to the larger firms already in the field, such as the accounting and auditing firm Accenture [11], and become a services line with an ordinary client relationship. Or a party other than the labs could pay for the ecosystem Venkatasubramanian describes [5].

We think the first is the hardest to reach. An evaluator sees only what lab staff are allowed to show it, so the lab's information rules limit the check whatever the contract says. OpenAI fired three employees last week for "violating our policies on accessing and handling sensitive company information," a spokesperson said. Two of them, Mikita Balesni and Tomek Korbak, said they believe they were dismissed over how they communicated with third-party evaluators [15]. "I worry the pervading fear to speak up and engage with third parties will mean OpenAI will cut corners on safety behind closed doors," Balesni wrote on X [16]. The counter-case is that OpenAI is fencing off sensitive data the way it would for any contractor, and that signed contracts will replace informal staff contact with something more durable. OpenAI disputed Balesni's characterization and wrote that it is "committed to embedding external assessors" [14].

We would be wrong if the contract details OpenAI has promised [8] give evaluators access and a reporting line that do not run through the company. Andrew Freedman, chief executive of the policy nonprofit Fathom, said the field is maturing quickly. "I've never seen an issue move so fast on so many different political spectrums," he told CNBC [17].

What to watch

  • Which assessors OpenAI signs: the small nonprofits it already works with, such as METR, or larger audit firms like Accenture.
  • Whether Anthropic, which did not comment to CNBC, names its embedded evaluators and says who pays them.
  • Whether the White House's voluntary accord adds any requirement on evaluator access or reporting beyond urging labs to partner with one.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence55
Adoption25
Hype gap+25
Incentives65
Confidence50
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    Anthropic and OpenAI are seeking support from a handful of small third-party groups like Model Evaluation and Threat Research (METR), Apollo Research and Transluce.

    ReportedSupportedSource: CNBCView cited source
  2. [2]

    The evaluators mostly operate as nonprofits; their primary role has been to assess AI model capabilities and risks and to call attention to instances where the technology behaves badly.

    ReportedSupportedSource: CNBCView cited source
  3. [3]

    The evaluators are still finding their footing in an industry where capital is flowing at historic levels and new models are rolling out faster than ever.

    ReportedSupportedSource: CNBCView cited source

Sources

1 independent publisher whose own reporting we read for this story.

  1. cnbc.com

    1 article · October 11, 2026

    AI’s quiet safety gatekeepers are stepping into the spotlight

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Loading related stories