Build5 publishers2 min readPublished
Anthropic's 950 Claude agents spent 210 million tokens to hand biologists 20 enzyme leads
Anthropic ran about 950 Claude agents for 21 hours on 210 million tokens to cut over 200,000 reverse transcriptases to 20 lab candidates. Human scientists kept the bench work, so lab capacity sets how fast those leads get tested.
The Engineer · Build desk

What happened
- Between the raw haul and the shortlist, the agents flagged about 3,500 candidate systems, according to Anthropic's account.
- One agent noticed repeating DNA beside an unusual reverse transcriptase gene, a pattern the researchers had not asked it to look for.
- Anthropic named the system ART, and its early lab experiments found the repeat array produces short RNAs whose function is still unknown.
- Earlier studies had already described the reverse transcriptase itself, so the claim rests on the repeats and a partner gene forming a larger system.
- Lucas Harrington, a Mammoth Biosciences co-founder, wrote on X that the method is decades old and that similar systems have been known since 2008.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- cost Each of the 20 finalists took about 10.5 million tokens of agent work before any lab time, and a team copying the design pays roughly that per lead.
- constraint Because humans run every experiment, adding agents or tokens raises the number of leads without raising how many of them the lab can test.
- decision Teams planning an agent search now have a worked split to copy or reject: per Dario Amodei, humans picked the search area and ran the wet lab, and agents handled the triage.
- contradiction Anthropic presents ART as a discovery, while Harrington argues the hard part is showing what a system does, and ART's function has not been shown yet.
Spread across every enzyme gathered, 210 million tokens comes to at most about 1,050 tokens per reverse transcriptase [1]. Spread across the flagged candidate systems, it is about 60,000 tokens each [2]. Per agent, the run averaged about 221,000 tokens [4]. The fleet used about 10 million tokens an hour [5].
Anthropic reported only totals for time and tokens [1]. The accounts give no per-stage split, no model version and no price.
Those averages describe Anthropic's workload. For them to transfer, the target has to be a sequence database big enough that a pass over every record is worth paying for, and a lab has to be ready to take the finalists. The figures are also Anthropic's description of its own system, in a preprint that has not completed peer review [11].
The design choice I would defend is the handoff. The agents wrote a human-readable report on each finalist [4]. Scientists read those reports, and the 3,500 flags stayed with the agents [3]. The physical experiments happen in the Bay Area lab Anthropic built after forming its life sciences group in spring 2026 [9].
The 2024 SAMPLE system closed the loop differently. It sent its protein designs to automated lab equipment and used the results to pick the next test, all aimed at better heat tolerance [12]. SAMPLE was aimed at a fixed target, while Anthropic's agents worked from an open brief [5].
A brief to look for interesting enzymes would fail any spec review I have sat in [5], and it can also surface a pattern nobody named in advance. Before escalating the ART lead, the agent checked the repeat spacing, compared the system with known biology and searched published research [6].
CRISPR researcher Feng Zhang said the finding is "intriguing and merits further investigation" after reviewing the preprint, The Neuron reported [15]. On this evidence, the new part is that agents ran the triage stage of a method genome miners already use [13]. Finding out what ART's short RNAs do is bench work [8], and humans do the bench work [9].
What to watch
- Peer review of the ART preprint, and whether outside groups can reproduce the agent search with their own models and data.
- Lab results showing what ART's short RNAs do, the gap Lucas Harrington says Anthropic has not closed.
- A per-stage token breakdown or compute bill from Anthropic, to show whether the first pass or the triage used most of the 210 million tokens.