Skip to content

Build1 publisher3 min readPublished Updated

Anthropic scores Claude as leading 26% of measured R&D with a prototype that uses Claude to judge

Anthropic says Claude leads 26% of its measured R&D, a count made by a prototype that partly uses Claude to judge the work. OpenAI's 3.1 agent workdays per human workday measures effort, and neither figure shows whether AI is speeding up AI research.

The Engineer · Build desk

Illustration accompanying Anthropic scores Claude as leading 26% of measured R&D with a prototype that uses Claude to judge

What happened

  • Anthropic defines "leads" as Claude completing most of a task from a high-level prompt while a human still supervises the work.
  • OpenAI says its systems have reached an "automated research intern" stage, finishing well-defined assignments that would take a skilled researcher several days.
  • A paper from 22 researchers and technology leaders, including Geoffrey Hinton, Yoshua Bengio, Jakub Pachocki, Jack Clark and Eric Horvitz, asks governments to watch automated AI research closely.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision Staffing or timeline plans built on the 26% should budget a supervising human for every Claude-led task, because Anthropic counts no measured work as fully autonomous.
  • constraint Checking either lab's figure independently requires internal data the labs control access to, so outside planners are working from the labs' own accounts.
  • constraint The two figures have no shared unit, one a share of graded work and the other an effort ratio, so they cannot be combined into one estimate of how much AI builds AI.

According to The Neuron's reading of the disclosure, Anthropic's scheme grades categories of work by how automated they have become [14]. Nothing sits in the highest grade. Anthropic says Claude is not fully autonomous on any measured portion of the work [4]. Going by the definitions, "leads" sits below full autonomy and above a looser grade where Claude completes large chunks under human direction [2][3]. Suppose the Claude-led work is counted inside the more-than-90% that has reached at least the large-chunks level. Then more than 64 percentage points of measured R&D sits in that middle grade [1]. The most common state of Anthropic's measured research is Claude doing large pieces while a person directs it [1].

Every one of those shares is a share of measured work [1]. For the 26% to describe Anthropic's research as a whole, the work that gets measured would have to look like the work that does not [1].

The classifier is the part I would push on at review. The prototype partly uses Claude to classify and judge the Claude-assisted work it counts [5]. For the 26% to carry over into anyone else's planning, Claude's labels would need to match human reviewers' labels on the same tasks, at an agreement rate someone had actually measured. I would not accept the arrangement in a test suite, where code under test writing its own pass criteria is the first thing a reviewer flags. Anthropic does deserve credit for publishing a precise definition of "leads" and stating the prototype status alongside the headline number [2][5].

OpenAI's figure measures something else. It counts effort: OpenAI tracks agent use, experiment volume, code production and task duration [13]. Taken as a share of volume, 3.1 agent workdays per human workday puts agents at about 76% of research-organization workdays [3]. For that ratio to say anything about speed, an agent workday would have to produce accepted research at something near the rate of a human one. OpenAI's own caveat on that point is the most useful sentence in either disclosure. The company says more code and more experiments do not necessarily mean AI itself is advancing faster, and that human researchers still set priorities, interpret results and decide what gets scaled or deployed [8]. OpenAI calls the measurements preliminary [9].

The 22-author paper is just as careful [10]. Its authors say the intelligence explosion they describe has not happened [16]. They also list what could break the loop: compute constraints, slow training runs, diminishing returns, difficult research problems, and tasks where humans remain the bottleneck [11]. The Neuron's assessment is that no independent evidence yet shows the labs' gains have created that feedback loop [15].

What to watch

  • Whether Anthropic publishes how often Claude's classifications in the R&D measurement prototype agree with human reviewers on the same tasks.
  • Whether OpenAI moves from effort measures such as agent workdays and experiment counts to output measures such as time to reach a given capability gain.
  • Whether any government answers the 22-author paper with a reporting requirement that gives outsiders access to the labs' internal automation data.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories