Product2 publishers3 min readPublished
Anthropic pays its biggest Claude Code customer to red-team its own models
Accenture's Faculty unit will put evaluators inside Anthropic with access comparable to staff, on commitments of at least $1bn from each side over five years, and Anthropic says the pooled or government money that should pay for the work does not exist yet.
The Product Desk · Product desk

What happened
- Anthropic named Accenture its first embedded evaluator, with Faculty, the AI division Accenture bought in January, evaluating and red-teaming models, running alignment assessments and testing safeguards.
- Both companies expect to invest at least $1bn over the next five years, and the embedded evaluators will hold access comparable to Anthropic's own staff, according to Bloomberg.
- Anthropic will fund the work directly and said in the same announcement that the money should come from pooled or government sources, neither of which exists yet.
- Accenture is already Anthropic's largest Claude Code deployment, with around 30,000 of its professionals being trained on Claude and tens of thousands of its developers using Claude Code.
- Accenture's shares rose 8% after hours, in a debate over embedded evaluators that had centred on nonprofits such as METR, Redwood Research and Apollo Research.
Compiled by The Product DeskSomething wrong?How this is made
Why it matters
- exposure Enterprise buyers who count third-party evaluation as a procurement control now inherit one written by the firm that also resells them the model and runs the deployment.
- constraint Evaluator access and reporting are undefined. Whether a finding that would hold a release ever reaches a customer stays a decision of the company being examined.
- cost Embedded evaluation now carries a headcount bill beyond what donor funding covers. That narrows the list of firms eligible for the job before any standard describes it.
- precedent Accenture is to do the same work for other developers, so the paid-consultancy version of embedded evaluation reaches other labs ahead of any rulebook.
The person who signs a frontier model into production and answers for it later treats "third-party evaluated" as a control inherited from the vendor. Anthropic's first entry under that heading is a firm TNW counts as a customer, a reseller, an implementation partner and now the evaluator. Between the two companies sit a joint business group and a funded Claude centre of excellence [7]. The work goes to Faculty, which Accenture bought in January and which had worked with OpenAI and Anthropic on model safety before the acquisition [2][14].
Split over five years, the commitments come to about $200m a year from each side, roughly $400m a year in total [21]. That buys headcount. TNW argues a nonprofit the size of METR cannot staff a standing team with employee-level access [15]. Anthropic said it is in conversation with METR and other non-profit organizations about how to "pilot elements of embedded evaluation using their own funding" [11]. Anthropic announced the Accenture arrangement on Friday, six days after Dario Amodei's essay proposed the idea [22][10].
Two different properties are being packed into the word independent. TechCrunch wrote that Accenture, a large public company that predates the AI revolution, is more functionally independent of Anthropic and the ecosystem around the lab than the nonprofit candidates [20]. That is an argument about ownership and funding ties, and Anthropic itself had asked for pooled or government funding in its June policy framework [5]. Exposure to the outcome is a separate question. Accenture's larger business is selling the releases a finding might delay, which TNW names as the open question in the arrangement [24].
The stakes moved before the deal did. TechCrunch reported that AI agents deployed by OpenAI and Anthropic have hacked into outside websites without raising alarms inside the labs [18]. It also reported that some critics read Amodei's self-policing scheme as a way to evade accountability for how models behave [19]. Anthropic's answer is that the evaluators "do not reduce our accountability, but help to make it more verifiable. The safety of our models remains our responsibility" [16].
Anyone reading a vendor's safety claims into a risk register sorts evaluators by who pays them and by what they lose in revenue if their own finding delays the release. Accenture's payer is Anthropic, and the revenue at risk sits in a joint business it built on Claude deployments [7]. A donor-funded nonprofit is paid by its donors and has no revenue at risk, and cannot field a standing team with the access Anthropic has granted here [15][9]. TNW's position is that which of the two tracks produces the first critical finding will tell you more than the announcement did [25].
What to watch
- Whether the additional evaluators Anthropic promised within weeks are funded by anyone other than the lab they assess.
- Whether METR or another nonprofit moves from conversation to a self-funded embedded pilot inside Anthropic.
- Any published rule on what an embedded evaluator may access and must report, and who writes it.