Product1 distinct publisher3 min readUpdated
Two contract labs made and tested Claude's designs, which is genuinely new. The field baseline those results are scored against comes from the company that ran the experiment.
The Product Desk · Product desk

Compiled by The Product DeskSomething wrong?How this is made
Anthropic published two experiments on Tuesday claiming its Claude models designed working protein binders and ran a routine chemical analysis in minutes, and it reported the results itself on its research blog [s1c1]. The part worth attention is not the hit rate: it is that the physical checks were done by two outside firms, Adaptyv Bio and Twist Bioscience, which independently produced and tested the designs, according to Anthropic [s1c5].
That matters because wet-lab validation is the step that software does not compress. Anthropic says Claude designed binders against 15 targets and succeeded against 14, with 22 to 35 percent of designs binding their target depending on the setup [s1c2][s1c3]. The campaign produced 354 confirmed binders from 1,320 designs [s1c10], which works out to about 26.8 percent overall [1]. The company says 10 to 15 percent is typical in the field today [s1c4]. That baseline is the load-bearing number in the whole story, and it is Anthropic's characterisation, not a cited independent measurement.
The one comparison that does have an external reference point is RBX1, where Anthropic says a preview of its Mythos model hit 40 percent against 3.7 percent for human entrants in an Adaptyv Bio competition, with Claude's top design beating the winning entry [s1c11]. That is roughly an eleven-fold gap [2]. It is also a comparison in which Adaptyv Bio appears twice: as the source of the human benchmark and as one of the two labs that made and tested Claude's designs [s1c5][s1c11].
The failures are more informative than the wins. Against maltose-binding protein, a notoriously smooth target, none of 90 designs was confirmed to bind [s1c13]. Results against BBF-14, a synthetic protein used as a hard benchmark, were only modest [s1c14]. On TNF-alpha, Opus 4.8 produced binders that worked across human, monkey and mouse versions while the stronger Mythos preview failed, and Anthropic said it does not know why [s1c12].
The compute was not modest either. One mode gave the models up to 12,500 Nvidia H100 hours across a 48-hour session [s1c8], implying on the order of 260 GPUs running in parallel [3]; a second mode used up to 2,500 H100 hours per target [s1c8]. The models were given a long prompt, internet access, connectors and a GPU pool inside Claude Science, then left to work autonomously, choosing binding sites and orchestrating existing open-source design and folding models [s1c6][s1c7].
The chemistry experiment is the cleaner result, because it has an answer key. Anthropic gave Claude Opus 5 raw NMR and LC-MS files from a contract lab and a two-sentence prompt, with no vendor software and no operator, and the model returned processed results in 23 and 19 minutes [s1c15][s1c16]. Purity came out at 96.4 percent against the lab's 96.33 percent [s1c17]. The model also caught and corrected its own overstatement of shifted peaks, and proposed the same confirmatory test the lab had run three days earlier [s1c18].
Anthropic says it plans further testing to confirm its hit rates and binding measurements, and has released its prompts and data [s1c19]. Watch whether anyone outside the company reruns the campaign against a published field baseline rather than a described one, and whether Adaptyv Bio or Twist Bioscience publish their own accounts of what they synthesised and measured. Until then the verification story stops at "the designs were real proteins," which is not the same as "they were better than what people do."
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Anthropic published two experiments on Tuesday on its research blog, saying its Claude models designed working protein binders and ran a chemical-analysis job in minutes; the company reported the results itself and framed them as early evidence Claude can speed up parts of drug development.
The second experiment used Claude Opus 5, the most capable model Anthropic offers to the general public, on chemical analysis: checking what a compound is and how pure it is, work chemists normally do with NMR and LC-MS.
Anthropic gave Opus 5 only the raw files from a contract lab and a two-sentence prompt, with no vendor software and no operator; working inside Claude Science, the model returned processed results in 23 and 19 minutes, running the two analyses in parallel, the company said.
The company said it plans further testing to confirm its hit rates and binding measurements, and has released its prompts and data.
Anthropic said Claude designed protein binders against 15 targets and succeeded against 14.
Between 22 and 35 percent of Claude's designs bound to their target, depending on the setup, according to Anthropic.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Self-reported results with partial external wet-lab validation
The strongest evidentiary element is real: two independent contract labs produced and tested the designs, prompts and data were released, and disclosed failures (0 of 90 on maltose-binding protein, modest BBF-14 results, an unexplained model inversion on TNFα) are the kind of detail a purely promotional disclosure omits. Against that, everything is reported by the vendor about its own models, has not been peer reviewed, and the comparison baseline is vendor-supplied; the cluster contains one publisher relaying that disclosure with no independent expert assessment.
Lab-stage demonstration, gated capability
Observable uptake is confined to the experiment itself: two contract labs synthesised and tested designs, and artifacts were published. There is no disclosed customer, partner pipeline or production use, the binder work ran on a preview model plus Opus 4.8 rather than a generally available one, and protein design and other sensitive biology remain blocked on Anthropic's most capable public model pending a vetted access program. The chemistry run used a publicly available model, which is the only part with a plausible near-term user path.
Modestly overstated by framing, not by the outlet
The capability claim outruns its verification in two specific ways: hit rates are scored against a field baseline the experimenter supplied, and the 'beat human experts' comparison rests on one target in one competition. Anthropic's stated ambition of running drug development end to end sits far ahead of a first-step designed binder that has not been peer reviewed. The gap is kept modest because the disclosure includes failures, released artifacts, external wet-lab testing and the company's own acknowledgement that more characterisation is needed, and because the single outlet in the cluster foregrounds the self-reporting problem rather than amplifying it.
Vendor reporting on its own models and workbench
Anthropic designed the experiment, chose the targets and modes, defined the comparison baseline, scored the outcome, and sells both the models and the Claude Science workbench the work ran inside. It also has a strategic interest in the safety narrative it attaches — dual-use risk, blocked capability, a forthcoming vetted access program — which positions it as gatekeeper of the same capability. Mitigating factors are outsourced wet-lab validation, published prompts and data, and disclosed failures.
Single publisher relaying a single vendor disclosure
The factual record is internally consistent and specifically quantified, and the outlet is careful about attribution, which supports moderate confidence in what was claimed. Confidence in what the claims mean is lower: one publisher, one primary source, no peer review, no independent expert commentary, and key quantities (baseline, affinity gains) verifiable only against Anthropic's own forthcoming characterisation.
leadership
The AI bill nobody reconciles: cost per finished task, not per million tokens1 distinct publisher
build
Claude Code's new default is a confession: the approval prompt was never a control1 distinct publisher
build
Developer habit, priced at $965B: what Anthropic's run actually proves1 distinct publisher
invest
GLM-5.3 Buys Buyers Time: Z.ai's Coding Model Cuts Tokens, Not the Closed-Model Lead1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 19, 2026