Product2 publishers3 min readPublished
Anthropic reports Claude leading 26% of its model R&D on a definition it wrote itself
Anthropic says Claude leads 26 percent of its model research and development and takes part in more than 90 percent of it. The company wants rival labs to publish the same index so the numbers can be compared over time.
The Product Desk · Product desk

What happened
- Anthropic said Thursday that Claude is helping develop the next, more intelligent version of the model, and that Claude now leads 26 percent of the company's model research and development.
- The company says AI does at least large chunks of work under close human direction on more than 90 percent of its research, and counts the 26 percent inside that figure.
- The share Claude leads was none in February and reached a quarter in August, six months later, and about 30,000 agents were doing research and engineering work as of August.
- Anthropic derived the number from the first of three proposed measurements, an index of research it performs itself scored against an automation rating scale developed by Epoch AI.
- Alongside the metrics, Anthropic said it has committed to external third-party evaluators embedded within the company to monitor its safety efforts.
Compiled by The Product DeskSomething wrong?How this is made
Why it matters
- precedent Anthropic has set the format for this kind of disclosure. A rival lab that publishes now either adopts Anthropic's index and Epoch AI's scale or has to argue for a different one.
- constraint Third-party validation covers the method, while the count itself rests on Anthropic's inventory of research tasks and its own definition of what leading one means.
- decision Anyone quoting 26 percent as a target for their own team has to write their own definition of leading first, because the figure describes lab work supervised by the people who built the model.
- exposure The oversight measures Anthropic proposed, including how much agent activity is monitored and how often behaviour is flagged, give a customer or a regulator a specific thing to request from any lab.
A researcher writes a paragraph of instruction, hits return, and reads what comes back. Under Anthropic's scale, that exchange is what counts as Claude leading: the model completes most of a task end-to-end from a high-level prompt while a human supervises [3]. The score comes off Epoch AI's automation rating scale, which grades how much of a task the model performed [9]. The disclosure does not say what the 26 percent is a share of. A reader cannot tell whether a quarter means a quarter of the hours, a quarter of the tickets, or a quarter of a sampled task list.
The 26 sits inside the 90. Anthropic counts the led work as part of the wider collaborative figure [6], so subtracting one from the other leaves about 64 points of research and development where Claude does large chunks and a human directs each one [19].
The Associated Press reported that the led share was none in February and reached a quarter in August, six months later [7]. That averages roughly 4.3 points a month if the climb were even [20]. Engadget describes the same chart as plotting Claude's automation level since August 2025 [14], so the two accounts put different windows on one series.
The disclosure landed while leading figures in AI, Anthropic chief executive Dario Amodei prominent among them, were calling for a slowdown over safety concerns [22]. Anthropic asked other developers to publish similar metrics regularly under a public methodology, so numbers could be compared over time and potentially across labs [12], and said any frontier model maker could reproduce the index from its own data with third-party validation [13]. "We should do everything possible to minimize the gap between what frontier labs know and what the public knows," the company said in a blog post [15]. It also said models accelerating their own development could make it "more challenging for humans to understand or control these systems" [16]. According to the AP, it was unclear from the disclosure how close Anthropic believes it is to recursive self-improvement, which the company describes as a model's ability to autonomously build its successor [17][21].
For anyone who has to report their own AI adoption to a leadership team, the definition is the reusable part. Two questions sort the work: whether the model started from a high-level prompt or from a specified subtask, and whether a human read the output before it shipped. High-level prompt with a human reading is Anthropic's "leads". Specified subtask with a human reading is the rest of its 90 percent. High-level prompt with nobody reading is the cell Anthropic says is empty, because Claude is "not operating fully autonomously for any measured subset of AI R&D work" [4].
What to watch
- Whether OpenAI or another frontier lab publishes a comparable AI-led R&D figure, and whether it uses Epoch AI's automation scale.
- Whether Anthropic's embedded third-party evaluators publish anything on the agent oversight numbers: share monitored, review time, flag rate.
- Whether the next update reports a led share above 26 percent, and whether the definition of "leads" is unchanged.