Product1 publisher3 min readPublished
Anthropic's protein binders got tested by outside labs. The benchmark is still Anthropic's.
Two contract labs made and tested Claude's designs, which is genuinely new. The field baseline those results are scored against comes from the company that ran the experiment.
The Product Desk · Product desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Anthropic published two experiments on Tuesday on its research blog, saying its Claude models designed working protein binders and ran a chemical-analysis job in minutes; the company reported the results itself and framed them as early evidence Claude can speed up parts of drug development.
- Anthropic said Claude designed protein binders against 15 targets and succeeded against 14.
- Between 22 and 35 percent of Claude's designs bound to their target, depending on the setup, according to Anthropic.
- Anthropic says 10 to 15 percent is a typical binder success rate in the field today.
- The wet-lab checks were not done in-house: two outside firms, Adaptyv Bio and Twist Bioscience, independently produced and tested Claude's designs, Anthropic said.
Compiled by The Product DeskSomething wrong?How this is made
Why it matters
Anthropic published two experiments on Tuesday claiming its Claude models designed working protein binders and ran a routine chemical analysis in minutes, and it reported the results itself on its research blog [s1c1]. The part worth attention is not the hit rate: it is that the physical checks were done by two outside firms, Adaptyv Bio and Twist Bioscience, which independently produced and tested the designs, according to Anthropic [s1c5].
That matters because wet-lab validation is the step that software does not compress. Anthropic says Claude designed binders against 15 targets and succeeded against 14, with 22 to 35 percent of designs binding their target depending on the setup [s1c2][s1c3]. The campaign produced 354 confirmed binders from 1,320 designs [s1c10], which works out to about 26.8 percent overall [1]. The company says 10 to 15 percent is typical in the field today [s1c4]. That baseline is the load-bearing number in the whole story, and it is Anthropic's characterisation, not a cited independent measurement.
The one comparison that does have an external reference point is RBX1, where Anthropic says a preview of its Mythos model hit 40 percent against 3.7 percent for human entrants in an Adaptyv Bio competition, with Claude's top design beating the winning entry [s1c11]. That is roughly an eleven-fold gap [2]. It is also a comparison in which Adaptyv Bio appears twice: as the source of the human benchmark and as one of the two labs that made and tested Claude's designs [s1c5][s1c11].
The failures are more informative than the wins. Against maltose-binding protein, a notoriously smooth target, none of 90 designs was confirmed to bind [s1c13]. Results against BBF-14, a synthetic protein used as a hard benchmark, were only modest [s1c14]. On TNF-alpha, Opus 4.8 produced binders that worked across human, monkey and mouse versions while the stronger Mythos preview failed, and Anthropic said it does not know why [s1c12].
The compute was not modest either. One mode gave the models up to 12,500 Nvidia H100 hours across a 48-hour session [s1c8], implying on the order of 260 GPUs running in parallel [3]; a second mode used up to 2,500 H100 hours per target [s1c8]. The models were given a long prompt, internet access, connectors and a GPU pool inside Claude Science, then left to work autonomously, choosing binding sites and orchestrating existing open-source design and folding models [s1c6][s1c7].
The chemistry experiment is the cleaner result, because it has an answer key. Anthropic gave Claude Opus 5 raw NMR and LC-MS files from a contract lab and a two-sentence prompt, with no vendor software and no operator, and the model returned processed results in 23 and 19 minutes [s1c15][s1c16]. Purity came out at 96.4 percent against the lab's 96.33 percent [s1c17]. The model also caught and corrected its own overstatement of shifted peaks, and proposed the same confirmatory test the lab had run three days earlier [s1c18].
Anthropic says it plans further testing to confirm its hit rates and binding measurements, and has released its prompts and data [s1c19]. Watch whether anyone outside the company reruns the campaign against a published field baseline rather than a described one, and whether Adaptyv Bio or Twist Bioscience publish their own accounts of what they synthesised and measured. Until then the verification story stops at "the designs were real proteins," which is not the same as "they were better than what people do."