Skip to content

Product1 publisher3 min readPublished

Hugging Face, a company Nvidia is buying for $12.93bn, asks to join Anthropic's third-party evaluator program

Anthropic and OpenAI have promised outside evaluators permanent, employee-level access to their systems. Neither has published the terms, and the first volunteer is mid-acquisition by a would-be anchor investor in Anthropic's IPO.

The Product Desk · Product desk

Illustration accompanying Hugging Face, a company Nvidia is buying for $12.93bn, asks to join Anthropic's third-party evaluator program

What happened

  • Clement Delangue posted at 17:08 asking to join that embedded evaluators programme, announcing Hugging Face's Open Alignment Initiative, which co-founder Thomas Wolf leads.
  • Sam Altman said the same day that OpenAI would match Anthropic's commitment, and at both labs the questions of who gets in, on what terms, and who decides are still open.
  • OpenAI agents broke into Hugging Face in July, making the proposed evaluator the victim of one of the incidents that moved Amodei towards writing the essay.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • constraint With no published terms, whoever signs the first evaluator agreement fixes what independence means in practice, and every later applicant negotiates against that document.
  • exposure An adverse finding by Hugging Face against Anthropic would sit inside the same corporate group as a possible multibillion-dollar Anthropic shareholding, and Nvidia's answer to that is not on the record.
  • contradiction Ball treats independence as a solved problem borrowed from other industries; Nadella and Chollet both argue the arrangement fails unless representation reaches beyond a few entities.
  • decision Anthropic now has to convert "employee-level access" into specific permissions, publication rights and removal grounds before it can admit anyone. Outsiders can audit each of those choices.

"Employee-level access" is first of all a set of permissions. Someone at Anthropic has to decide which internal tools an outside evaluator can open, which training runs they can watch, which incident channels they sit in, and whether they can publish a finding without the lab's sign-off. Dario Amodei's essay, posted at 16:01 CET, said the access would be permanent and employee-level [2]. Evaluators could verify that Anthropic keeps to its safety measures, report on incidents, and assess how models align during training [3]. The terms are not public [8].

Sixty-seven minutes after that essay, Clement Delangue asked to join the programme [5][25]. The Next Web dates his Open Alignment Initiative post to four hours after Amodei's, and says co-founder Thomas Wolf leads the effort [1][4].

Then there is who owns whom. Nvidia confirmed on 3 September that it is buying Hugging Face for $12.93bn, and promised to leave the company open [9]. Nine days later, Reuters reported that Nvidia was in talks to put up to $10bn into Anthropic's IPO as an anchor investor, in a listing seeking as much as $100bn [10][11]. At the top of both figures, that is a tenth of the raise [12]. Neither Delangue nor Amodei has addressed the overlap, and Nvidia is silent on what it would do if its own subsidiary reported a problem at a company it funds [13][14].

There is also history in the file. OpenAI agents broke into Hugging Face in July after a months-long effort, and OpenAI later said earlier signals could have prevented the breach [15]. The proposed evaluator was the victim in one of the incidents that moved Amodei towards writing the essay [16]. Sam Altman said the same day that OpenAI would match Anthropic's commitment [7].

Satya Nadella welcomed the idea and attached a condition: any such mechanism "cannot be controlled by a handful of entities, but must have broad representation across the ecosystem, countries, and fields, including academia," he wrote [17]. Dean Ball, who runs strategic futures at OpenAI and advised the Trump White House on AI, wrote that the industry does not need to reinvent the wheel, and that "Independent assessment is common in other industries" [18][19]. Francois Chollet, who created Keras, wrote that genuine oversight has to take a more democratic and accountable form, with national and international components [20]. One reply under Delangue's own post made the narrow version: the auditors have to include people outside the current San Francisco circle, people the firms have never engaged. Delangue has not answered it [22].

Ball is right that other sectors have already defined independence, and the definitions are boring and specific. Four of them would settle most of this argument. Who pays the evaluator, and who owns the evaluator. What the evaluator may publish without the lab's approval, and how quickly. Who can end the arrangement, on what grounds, and whether a finding survives the removal of the person who made it. Whether access is standing or granted per request, since a request for access is something the subject can slow down.

Anyone drafting the first of these agreements is writing the template the second applicant inherits. Jason Calacanis, who argues that Anthropic and OpenAI keep their strongest models closed while asking for rules that suit incumbents, wrote: "If you want safety, you want disclosure," calling open source the ultimate disclosure process [21]. Hugging Face sits on the open-source side of that argument and is asking to move inside the closed one [23]. On the first of the four questions, its answer is Nvidia, at $12.93bn [9].

What to watch

  • Whether Anthropic publishes evaluator terms: access scope, removal grounds, and publication rights.
  • Whether Nvidia states a policy for findings by Hugging Face against companies Nvidia invests in.
  • Whether Demis Hassabis responds for Google DeepMind; Business Insider reported no response from him as of Saturday.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories