Build1 publisher2 min readPublished
Anthropic's agents filtered 200,000 DNA sequences to one, almost all in software
Anthropic says roughly 950 Claude agents spent 21 hours and 210 million tokens narrowing 200,000 DNA sequences to one new enzyme system named ART. The filtering ran almost entirely in software, the part of the setup builders can copy.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Claude narrowed the sequences to 3,500 candidate systems, then wrote full human-readable reports on only the 20 it judged most compelling.
- Anthropic says human work was limited to the initial prompt and the lab bench, with the agents combing the database and judging candidates on their own.
- Anthropic says the same kind of analysis would take an expert scientist weeks to months of work.
- The Hacker News thread on the announcement stood at 567 points and 588 comments, with most of the argument over whether 'discovery' is the right word.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- capability A small team can now run an exhaustive database survey as an orchestration job, getting back a short list of candidates to verify instead of doing the survey by hand.
- constraint The method only works where a survivor can be verified. Here the oracle is a wet lab; a software team needs a test suite or equivalent, or the self-review has nothing to calibrate against.
- constraint Whether a copy reproduces the result turns on the retrieval step as much as the model, because detection collapses once the agent has to locate the data itself.
- cost The payoff rests on selectivity: software here spared roughly five candidates in a million, so the parallel token spend only pays when survivors are that rare.
"Fan out cheap, filter hard, escalate survivors." That is how the dev.to teardown describes the shape. [18] Anthropic says the agents found a new enzyme system hiding in phage DNA, in a post dated September 23, and the search ran over one large database of DNA sequences. [1][5] What differs from a normal agent run is the ground-truth check. For software it is usually a test suite; here it was a wet lab. Almost everything died in software, before any bench time. [18]
The coordination was custom. The tools were Claude Science and Claude Code, which Anthropic calls "the same tools available to any scientist," plus a harness of its own "that coordinates many Claude sessions running in parallel." [11][12] The post does not name a specific model. [13]
The pipeline front-loads a calibration step. Before hunting anything new, Anthropic says, Claude "reads the relevant literature and reproduces the established results from public data to check its methods." [15] Then it looks for genomic neighbors that fit no described system and runs an adversarial self-review. "Typically most candidates are eliminated at this stage," Anthropic writes. "A survey may end with a single candidate worth testing, or with none." [16]
The agent that found ART left a trace a human could check. It "counted the repeats and measured their spacing, compared the layout with the known RT systems, and searched the literature for any previous report of the pattern" before filing its report for review. [17] The teardown's author learned to value that trail the hard way, after an agent they built reported a prompt injection that had never happened and the logs said otherwise. [23]
The reproducibility figures are the catch for anyone copying the pattern. The Next Web's account of the pre-print says reruns of the same search missed the repeat array in all 10 attempts. [19] Detection ran around 90% when the DNA region was handed to the model directly, and as low as 32% when the model had to find it through the tool loop, roughly a third of the direct rate. [20][21][2]
What to watch
- Whether other labs reproduce the ART system from the same public data, and whether the pre-print clears peer review.
- Whether Anthropic releases the harness or names the model, so the pipeline can be rebuilt rather than inferred from a blog post.
- Whether the tool-loop detection gap narrows as teams tune the retrieval step feeding the agents.