Published · 4d agoBuild3 min read
Anthropic's protein run is checkable, which is rarer than the 26.8% hit rate
Claude agents designed 1,440 protein binders and 354 bound. The interesting part is that the prompts, provenance and every measurement went out with the number.
Written for builders.See today for builders

What happened
- Anthropic described the protein design results in a post on X on Tuesday and published a 29-page technical report.
- Anthropic released the prompts, predicted structures, provenance records and experimental measurements for all 1,440 designs on Hugging Face.
- Amir Shanehsazzadeh, an Anthropic researcher and the report's sole listed author, wrote a roughly 16,000-word protocol prompt covering the scientific and operational knowledge needed to run a binder campaign; Claude handled the individual design decisions.
- Anthropic ran Claude Opus 4.8 and Mythos Preview as agents against 16 protein targets. The models researched each target, selected regions and binding sites, installed open-source design and structure-prediction software, generated candidates, optimized them computationally and returned 30 ranked designs per target.
- Claude received no predetermined epitope, scaffold or protein sequence.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
Anthropic described a set of autonomous protein binder campaigns in a post on X on Tuesday and published a 29-page technical report [1], then released the prompts, predicted structures, provenance records and experimental measurements for all 1,440 designs on Hugging Face [2]. That second act is the one operators should care about: a hit rate is a marketing artifact until someone outside the lab can recount it, and this one can be recounted.
The setup is legible. Amir Shanehsazzadeh, the report's sole listed author, wrote a roughly 16,000-word protocol prompt covering the scientific and operational knowledge needed to run a binder campaign, and Claude made the individual design decisions [3]. Anthropic ran Claude Opus 4.8 and Mythos Preview as agents against 16 targets; the models researched each target, picked regions and binding sites, installed open-source design and structure-prediction software, generated and computationally optimized candidates, and returned 30 ranked designs per target [4]. No epitope, scaffold or sequence was supplied in advance [5]. Single-target runs got 24 hours and $10,000 of compute; multi-target runs got 48 hours and $50,000 [6]. Anthropic chose the targets, supplied the GPU accounts, set the budgets, ordered synthesis and interpreted the results afterwards [7]. Notably, Anthropic introduced no specialized protein-generation model; Claude coordinated existing open-source backbone generators, sequence-design tools and structure predictors [8].
Adaptyv Bio and Twist Bioscience, both paid contract research organizations, synthesized and measured the designs [9]. One target, mature GDF-8, aggregated during testing, leaving 1,320 designs across 15 targets with interpretable measurements [10]. Of those, 354 bound, a reported 26.8% [11], which recomputes to 26.8% [1]. Claude produced at least one binder against 14 of the 15 interpretable targets [12], and 49% of designs ranked first in their campaign bound [13].
The averages hide most of the signal. TREM2 returned 72 binders from 90 designs and VEGF-A 54 from 90 [14], or 80% and 60% [2]. The synthetic beta-barrel BBF-14 produced three weak binders, and maltose-binding protein produced none from 90 designs [15]. Anthropic reports that the computational confidence scores gave little warning those campaigns would fail [16], which is the finding a buyer should weigh most heavily.
Here is where the released files earn their keep. The summary numbers do not fully reconcile on their own: 1,440 designs minus 1,320 interpretable leaves 120 excluded for a single target [3], while the per-target counts quoted are 90 [4], and 15 targets at 90 designs each would be 1,350, or 30 more than the stated 1,320 [5]. With the dataset published, that is an arithmetic question anyone can settle. Without it, it would be a footnote nobody could audit.
The model comparison is similarly hedged in the source itself. Mythos Preview's single-target campaigns hit 35.1%, against 26.7% for its multi-target run and 22.6% for Opus 4.8 multi-target [17], a spread of 12.5 points [6]. The single-target format also received 2.8 times the compute per target, and Anthropic says the experiment cannot isolate whether focus, budget or model behavior caused the gap [18]. The sharpest external comparison is RBX1, where 28 of Claude's 90 designs bound against nine of 245 entries in a recent open competition [19]: 31.1% versus 3.7% [7].
Watch whether anyone outside Anthropic recounts the Hugging Face measurements and publishes a diff against the report. Watch for a compute-matched single-versus-multi-target run that answers the question this one cannot. And watch whether the maltose-binding protein and BBF-14 failures get a mechanistic explanation, since confidence scores did not predict them [16]. Anthropic also opened Claude Cowork to every paid plan earlier the same Tuesday, according to RuntimeWire [20]; the orchestration pattern is the product, and the protein work is its most auditable demonstration so far.
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
Anthropic described the protein design results in a post on X on Tuesday and published a 29-page technical report.
ReportedView cited source - [2]
Anthropic released the prompts, predicted structures, provenance records and experimental measurements for all 1,440 designs on Hugging Face.
ReportedView cited source - [3]
Amir Shanehsazzadeh, an Anthropic researcher and the report's sole listed author, wrote a roughly 16,000-word protocol prompt covering the scientific and operational knowledge needed to run a binder campaign; Claude handled the individual design decisions.
ReportedView cited source - [4]
Anthropic ran Claude Opus 4.8 and Mythos Preview as agents against 16 protein targets. The models researched each target, selected regions and binding sites, installed open-source design and structure-prediction software, generated candidates, optimized them computationally and returned 30 ranked designs per target.
ReportedView cited source - [6]
Single-target runs received 24 hours and a $10,000 compute budget; multi-target runs received 48 hours and $50,000.
ReportedView cited source
Sources & coverage · 2 publishers
The reporting this story was synthesized from, earliest first. Every link goes to the original.
- runtimewire.comRyan Merket4d agoAnthropic says 354 Claude protein designs bound in lab tests
- the-decoder.comMaximilian Schreiner4d agoAnthropic says any lab can now let a language model agent run the whole protein design stack
Additional citations
- RuntimeWire

