Leadership2 publishers3 min readPublished
Claude graded the index that credits it with leading 26% of Anthropic's R&D
Anthropic has put a number on how much of its own AI research Claude now runs. The number comes out of a pipeline in which a Claude agent catalogued the tasks and a separate Claude judge scored them.
The Board Room · Leadership desk

What happened
- Anthropic said Claude "leads" 26% of its AI research and development work as of August 2026, measured on an automation scale developed by Epoch AI that runs from no AI involvement to no human in the loop.
- The ratings came from Anthropic's own model: a Claude research agent assessed each category and a separate Claude judge assigned the automation level, with employees scoring their own work areas blind to the model's evidence.
- For the week of July 13 to 20, Anthropic classified about 6% of its AI R&D compute as safety work, rising to about 12% of the compute used in AI-driven AI R&D.
Compiled by The Board RoomSomething wrong?How this is made
Why it matters
- constraint The 542-node basket was frozen in July, so every later reading scores a fixed task list. An index that climbs on that basket cannot tell a manager whether the underlying work changed shape.
- decision AL4 keeps the deployment call with an engineer, so the number is not evidence that senior sign-off has moved. A leader citing 26% to justify headcount decisions is extrapolating well past what was measured.
- exposure The figures are Anthropic's own, produced largely by Claude, and the report does not say any outside party has checked them. Anyone quoting 26% to a board or a regulator inherits that provenance.
- contradiction implicator.ai dates the sub-1% baseline to February 2026 and Business Insider to March, so the slope everyone will quote is uncertain by a month at the point where it is steepest.
Anthropic published two agreement rates, and the distance between them is what a reader has to judge. Its model's ratings matched staff ratings exactly 59% of the time; pairs of employees matched each other 35% of the time [5]. Model and human ratings landed within one level of each other 97% of the time [6]. One reading is that the model applies the scale more consistently than the people whose work it scores. The other is that 59% is about 1.7 times a rate that is itself barely one in three [7], which says the scale is hard to apply. Anthropic acknowledged that "the 'judge' model could make the same kinds of errors as the model it is checking" [8].
The basket being scored was built in July and then locked. For each week of July 2026, Anthropic randomly sampled 20% of staff in the departments that make up its model R&D loop, and a Claude research agent reviewed each sampled person's week using Slack and internal documentation [9]. That produced about 15,000 granular tasks, which Claude arranged into 542 nodes, roughly 28 tasks per node [10][11]. Anthropic froze the basket for later measurements [10]. So the August reading describes July's work.
AL4 describes supervised work. Anthropic's own example is a failed data pipeline: Claude inspects the logs, finds and tests a fix, handles surprises and compares the repaired output with the last good run, and an engineer still decides whether to deploy it [14]. "Claude is not operating fully autonomously for any measured subset of AI R&D work," Anthropic said [13]. The company also said the share of work at or above "AI collaborates" is above 90% [15], which puts more than 64% of the measured work at collaboration levels below "leads" [16].
The oversight numbers are large in one direction and small in the other. About 30,000 agents ran at once on Anthropic's most-used internal research and engineering platform in August, every action passing through an online monitor, typically within seconds [17]. Blocking about one in 47,000 of more than a billion decisions works out to roughly 21,000 blocked actions in the month, each reviewed by a person within a week [18][19]. A separate offline system flags about 100,000 transcripts a week, and about 50 high-priority cases reach people weekly, near 215 across a 30-day month [20][21]. Anthropic wrote, "We have been doing this kind of monitoring for months, not years" [22].
A skeptic would say a lab grading its own automation with its own model will arrive at a flattering number. Two things in the report cut against that. Anthropic classified about 6% of its AI R&D compute as safety work in the week of July 13 to 20, and about 12% of the compute used in AI-driven AI R&D, calling both estimates conservative because mixed-purpose work counted as capabilities research [23]. It also said one week is enough to show the measurement can be made and not enough to show a trend [24]. Anthropic said it plans to embed third-party evaluators and is setting them up [12].
The surrounding argument shapes how the figure travels. An Anthropic researcher resigned last week and accused leading AI companies of "gambling with our lives", and chief executive Dario Amodei has since called for frontier AI companies to coordinate on slowing development, according to Business Insider [25]. Anthropic said other frontier labs could publish the same figures and allow third parties to verify them [26]. "We are reporting these measurements because they give the public, third parties, and governments better visibility into the pace of AI development inside frontier labs," the company said [27].
What to watch
- A second reading against the same 542-node basket, which would show whether 26% is a level or a slope.
- Whether third-party evaluators score the same nodes and land at lower automation levels than Claude's judge did.
- Any redefinition of the 542 nodes, which would reset the series before it has two comparable points.