Skip to content

Invest4 publishers3 min readPublished Updated

Anthropic invites its rivals to benchmark against a self-reported 26%

Anthropic says Claude now leads 26% of its model research and development under human supervision, and it wants other labs to publish comparable figures. The labs count different things.

The Investor · Invest desk

Photograph accompanying Anthropic invites its rivals to benchmark against a self-reported 26%
Photo: fortune.com

What happened

  • The disclosure gave rival labs a look at Anthropic's progress toward recursive self-improvement, and the company encouraged those competitors to publish similar metrics of their own.
  • OpenAI said this month it had built an automated "research intern" that handles well-defined research tasks under human direction, including tasks that would take a skilled researcher a few days.
  • Fortune reports that leading AI companies define recursive self-improvement differently, some counting any AI feedback on model improvement and others only fully autonomous self-design.

Compiled by The InvestorSomething wrong?How this is made

Why it matters

  • contradiction Anthropic wants comparable numbers from labs that do not agree on what to count, so a 26% from one and a research intern from another cannot be put in the same column.
  • decision Every rival lab now chooses between publishing in Anthropic's unit, publishing in its own, or staying quiet, and the second publisher effectively ratifies or rejects Anthropic's definition.
  • precedent OpenAI has attached a calendar date, March 2028, to an automated researcher. A date is something outsiders can hold a lab to.
  • constraint The missing denominator leaves two readings open: work taken off researchers' desks, or work reclassified into a category Claude already handles.

"Leading" 26% of model research and development means, in Anthropic's own gloss, that Claude can complete most of a given task end-to-end from a high-level prompt while a human still supervises [2]. So the figure sorts tasks by how much prompting they need, and it leaves 74% of that research and development not led by Claude [20]. Fortune's account explains what leading means; it does not report how Anthropic defines the total the 26% is a share of [21].

The invitation to competitors is hard to accept on its own terms [4]. Fortune reports that the leading labs do not agree on what recursive self-improvement is, with some counting any feedback from AI on model improvement and others only a system working toward it fully autonomously [6]. Under the loose version, the number would have been sizeable years ago. John Thickstun, a Cornell assistant professor of computer science who studies methods that control model behaviour, said: "We have already, for years, been using these models in supportive roles for creating the next version of these models. So people use the past generation of models to write code for the AI systems that then create the next generation" [7]. He said the older attempts, including those by OpenAI co-founder Andrej Karpathy, produced minor improvements and not big creative leaps [8].

OpenAI answered in a different unit. It said this month it had built an automated "research intern" that carries out well-defined research tasks under human direction, including "tasks that would take a skilled researcher a few days" [9], and it has put a date on the full automated "researcher": March 2028 [10], about 18 months out [19]. It also wrote in that announcement that "Whether and how to proceed must depend on our ability to preserve human control and on informed democratic choices about the benefits and risks" [11]. Elon Musk said in March that for xAI's Grok models, "humans are gradually getting less and less in the loop" [13].

In my view the disclosure practice is the more interesting part of this, or rather, the sequencing is. Anthropic published a figure and asked rivals to publish theirs [4], and whoever goes second either adopts Anthropic's definition of a task or explains why it prefers its own. Anthony Aguirre, president and CEO of the Future of Life Institute and a physics professor at UC Santa Cruz, said "You can see in these plots from Anthropic over time, more and more of research is being done by the AI and it's becoming closer and closer to fully autonomous" [14], and "I think this is probably the worst idea in the history of humanity to do this. And yes, they're doing it" [15].

What would settle it: a second Anthropic disclosure that pairs the percentage with a denominator and a research headcount or spend line that has not grown, which would make 26% a productivity figure with money behind it. If the percentage climbs while the definition moves, it is positioning. Anthropic itself has stayed quiet on how close it is to fully autonomous model improvement [5].

What to watch

  • Whether any competing lab publishes a figure in Anthropic's unit, and whether it supplies the denominator Anthropic left out.
  • Whether Anthropic's next update moves the percentage, the definition of a task, or both.
  • Whether OpenAI restates its March 2028 automated researcher target as a percentage of work completed, or lets the date slip.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories