Skip to content

Build3 publishers3 min readPublished

Anthropic's 26% figure grades Claude's R&D tasks by how much human steering they took

The lab says Claude leads 26% of its model research and collaborates on about 90%, with both tiers defined by how closely a human directs each task, and it wants rival labs publishing the same measure.

The Engineer · Build desk

Photograph accompanying Anthropic's 26% figure grades Claude's R&D tasks by how much human steering they took
Photo: yahoo.com

What happened

  • About 90% of the company's research and development is now done in collaboration with Claude, according to the same announcement.
  • The share Claude leads was none in February and had reached a quarter of the work by August, on the company's own account of its progress.
  • The company framed the metrics as a way to gauge how close leading labs are to recursive self-improvement, a model's ability to autonomously build its successor.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision Review capacity becomes the staffing variable for anyone copying this setup: work in the collaborative band is work a human has to direct closely, so each added agent buys supervisor hours as well as throughput.
  • constraint An engineering manager cannot reproduce this number for their own team without first inventing a task taxonomy and a classification rule, which makes cross-company comparison a matter of trust for now.
  • precedent By asking other developers to publish the same measures on a shared public methodology, Anthropic sets a disclosure format that rivals will now either adopt or decline in public.
  • exposure Embedding external third-party evaluators inside the company gives people outside it standing to check safety work from within, and makes those evaluators the first outside readers of numbers like these.

A task Claude "leads" is one it can complete most of "end-to-end from a high-level prompt" while a human supervises, in Anthropic's wording [3]. A collaborative task is one where the model does "large chunks of work under close human direction" [5]. The difference between the two tiers is how much steering each task needed, and the company said the model is not yet working completely autonomously [4]. The gap between the two shares is 64 percentage points of research and development where a person is directing the work closely [20].

For the 26% to transfer to another engineering organisation, you would need the task list it divides and the rule that sorts each task into a tier. The Economic Times account of the announcement does not report either, nor a count of the humans supervising the roughly 30,000 agents the company said were doing research and engineering work in August [7][22]. The growth number has the same dependency. A rise from none to 26% between February and August works out to about 4.3 points a month [21].

The other figure people quote alongside this one counts a different unit. The Sequence, opening a series on recursive self-improvement, wrote that Anthropic published an essay in June 2026 stating that as of May, Claude authored more than 80 percent of the code merged into its production codebase [16]. Merged lines and research tasks are not the same denominator, so the two percentages do not stack. The same newsletter recorded OpenAI's disclosure that GPT-5.3-Codex helped debug its own training process and manage parts of its own deployment, and DeepMind's AlphaEvolve turning up algorithmic improvements that land in the infrastructure other models train on [17][18]. "The question is which parts of the job it took, how well it does them, and what checks the work," The Sequence wrote [19].

Until recently there was nothing to instrument. The Sequence noted that for about sixty years, arguing about recursive self-improvement meant arguing in the abstract, because there was no system to point at [23].

Anthropic said that models accelerating their own development could make it "more challenging for humans to understand or control these systems" [10]. Its blog post located the remedy in disclosure: "This means better measuring the development of AI, reporting on it publicly, and giving society an opportunity to decide how to use this information," the company said [9]. It also said oversight measures matter for seeing how often monitoring systems detect agent misbehavior [12].

An Anthropic researcher resigned last week with a warning about the threats the technology poses to humanity [14]. Chief executive Dario Amodei, OpenAI's Sam Altman, Elon Musk and other tech leaders have since supported slowing development, while other tech leaders and President Donald Trump have pushed back [15].

What to watch

  • Whether OpenAI or Google DeepMind publish a comparable "led" percentage with a methodology an outsider can check.
  • Whether Anthropic releases the task taxonomy and the human headcount behind the 26% and 90% figures.
  • Whether the embedded third-party evaluators report how often monitoring systems detect agent misbehavior.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories