Leadership1 publisher3 min readPublished
Anthropic's own index rates Claude as leading 26% of its AI R&D work
The lab published a prototype index of how much of its own AI research Claude now leads, alongside a promise to embed outside evaluators with access comparable to what its internal risk teams get. The figures are dated August 2026.
The Board Room · Leadership desk

What happened
- Anthropic published measurement tools covering how much of its AI R&D is done by AI, how well it can oversee and intervene in agent actions on its systems, and how it allocates the compute behind new models.
- Its prototype R&D Automation Index, scored on a scale developed by Epoch AI, rates Claude as leading 26% of the company's AI R&D work as of August 2026.
- The company says it will embed independent third-party evaluators from multiple organizations and give them access to internal processes, systems and data comparable to its internal risk assessment teams.
Compiled by The Board RoomSomething wrong?How this is made
Why it matters
- constraint By Anthropic's own account the series can only be tracked against its earlier self, so a buyer cannot use the 26% to rank vendors until a common methodology exists.
- precedent Once one lab offers evaluators the access its internal risk teams have, that becomes the level a regulator or a large customer can ask a rival to match, and declining becomes a visible choice.
- exposure Anthropic has put its own pace on the record using a figure it says would move under a pacing agreement, so it now owns the job of explaining any rise in later snapshots.
- decision A government drafting transparency rules can adopt a metric set authored by a company it would regulate, which is quicker, or spend the time defining its own.
Two figures in Anthropic's snapshot sit next to each other and should be read together. The company puts 26% of its AI R&D at AL4, the level where Claude completes most of a task end-to-end from a high-level prompt while a human supervises [6][5]. More than 90% of the work is at or above AL3, where the model does large chunks under close human direction [7][5]. Subtract the first from the second and at least 64 percentage points of the catalogued work sits in that supervised middle [15]. Under a tenth of it involves less AI than that [16].
The index is a prototype. Anthropic built it by cataloguing every kind of AI R&D work done at the company, rating how automated each task currently is, and aggregating those ratings [3]. The scale comes from Epoch AI and runs from AL0, no AI involvement, to AL5, where AI operates fully autonomously with no human in the loop [4]. No measured subset of the work is at AL5 [8]. Anthropic says the reason to track the series at all is that models accelerating their own development could make them harder for humans to understand or control, and that the public needs a way to see how close the world is to a model fully autonomously building its successor [19].
The access commitment is the part an outsider can eventually test. Anthropic says it will embed independent third-party evaluators from multiple organizations and give them access to internal processes, systems and data comparable to what its internal risk assessment teams have, to verify safety practices, report incidents and monitor metrics like these [2]. The post does not name the organizations or say when they arrive [17].
Anthropic designed the index, scored itself using its own models and picked the quarter to publish. It names two obstacles to comparing this kind of reporting across labs: the lack of a common methodology, and its use of its own models to evaluate its systems [10].
"AI systems are becoming exponentially more powerful and have begun to automate more of the process of building themselves," Anthropic wrote [12]. Capability evidence is published on a separate track, through Responsible Scaling Policy risk reports that include evidence on how much its models are accelerating AI R&D [13]. The company also says it would expect these numbers to shift if there were coordination on pacing the frontier, of the kind its chief executive Dario Amodei has called for [11].
For a buyer this quarter, the index changes nothing in a contract, but it does supply a format. Anthropic says any frontier developer could publish these measures regularly using a public methodology, and that doing so would let the numbers be compared over time and potentially across labs [9]. A vendor asked at the next renewal for its own automation distribution now has a published shape to answer in.
What to watch
- Which organizations get embedded at Anthropic, on what terms, and whether the access holds when a finding is unflattering.
- Whether a second frontier developer publishes an automation distribution on a methodology an outsider can check.
- Where the 26% sits in the next snapshot, and whether Anthropic keeps publishing on a fixed cadence.