Science1 distinct publisher3 min readUpdated
AIntibody tested 511 antibodies from 29 organizations on three tasks. Ranking clones inside HCDR3 clusters was worse than random for every model but one.
The Scientist · Science desk

Compiled by The ScientistSomething wrong?How this is made
A consortium has published what it describes as a prospective, blinded benchmark of computational antibody design, testing 511 AI-designed or predicted antibodies from 29 organizations against experimental affinity and developability readouts [1]. The result matters for procurement more than for research: performance tracked the task, not the vendor, and two of the three tasks exposed failure modes that no retrospective evaluation would have shown [5][6][7].
The challenge, called AIntibody and explicitly modeled on CASP, set three problems: in silico affinity maturation from phase 1 sequencing outputs, affinity ranking within heavy-chain CDR3 (HCDR3) clusters of a selection output, and CDR design of proteins that were not in a selection output [2][3]. Affinity-matured antibodies were modeled effectively [5]. Ranking high-affinity clones inside clustered HCDR3 datasets was, with the exception of a single model, worse than picking clones at random [6]. Out-of-library design was highly variable across most submissions, with many failing to outperform standard selection campaigns [7].
Several groups did produce developable antibodies with affinities below 100 pM [4]. According to the organizers, those successes were exceptions that did not transfer across tasks [8], which is the finding with the most direct commercial consequence. Only one of the three tasks yielded broadly reliable performance [14]. Across the field, 511 submitted antibodies from 29 organizations works out to roughly 18 per group [13], so this is a survey of many methods at shallow depth rather than a deep audit of a few.
The reason task specificity bites here is that antibody engineering is multiparametric. Beyond affinity and kinetics, developability covers expression, melting temperature, polyreactivity and aggregation, and the difficulty is improving one property without degrading others [9]. The authors put the asymmetry plainly: there are many solutions to improve a specific antibody, and many more ways to destroy it [10]. A model that ranks well inside a matured lineage is not thereby a model that can rank clones in a selection output.
The organizers are candid about the benchmark's own limits. This first iteration evaluated a single antigen, was organized by the same consortium now reporting the results, and relied on organizer integrity rather than informatic safeguards for blinding [11]. That is a weaker structure than CASP, which evaluated hundreds of diverse protein targets across multiple rounds under independent governance [12]. It is still stronger than the status quo, which the same authors characterize as performance claims resting on retrospective analyses without additional experiments, increasingly in preprints and technical reports [15].
The operating implication is to scope pilots by task with a defined baseline. For a ranking task, the baseline is random clone picking; for de novo CDR design, it is your existing selection campaign [6][7]. A platform that clears one bar tells you nothing about the other.
What to watch: whether future AIntibody rounds add antigens, independent governance and informatic blinding, which the organizers say they intend to address [16]; whether the sequencing datasets released with the study, which the authors offer for comparative benchmarking, get used by groups that did not compete [17]; and whether the named gaps in affinity prediction, library-inspired design and cross-task generalization close in a second round [18]. Related efforts, including the Adaptyv EGFR binder competition, which is not restricted to antibodies, provide a second reference point [19].
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
The AIntibody challenge tested 511 artificial intelligence-designed or predicted antibodies from 29 organizations, validated with diverse experimental assays and anchored to experimentally measured affinity and developability, in a prospective, blinded benchmark.
AIntibody tested three tasks: in silico affinity maturation from phase 1 sequencing outputs, affinity ranking within heavy-chain complementarity-determining region 3 (HCDR3) clusters of a selection output, and CDR design of proteins not included in a selection output.
AIntibody's first iteration evaluated a single antigen, was organized by the same consortium now reporting the results, and relied on organizer integrity rather than informatic safeguards for blinding.
CASP evaluated hundreds of diverse protein targets across multiple rounds with independent governance.
AIntibody was a challenge inspired by the Critical Assessment of Structure Prediction (CASP), and represents a first step toward CASP-caliber benchmarking in antibody discovery.
Several groups in the challenge produced developable antibodies with affinities below 100 pM.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Prospective blinded wet-lab validation, single antigen and single round
The result rests on designs synthesized as full IgGs and measured under uniform protocols by independent laboratories, with affinity and a five-assay developability panel as readouts, which is far stronger than retrospective scoring. It is held back from a higher score because the first iteration covered a single antigen in one round, and the authors acknowledge blinding depended on organizer integrity rather than informatic safeguards.
Broad field participation, no production or clinical use shown
Twenty-nine organizations submitting 511 designs is a substantive participation signal that the method community engaged with a blinded evaluation, and the release of the underlying sequencing datasets supports further use. Adoption stays low because the sources show only benchmark participation and dataset provision — no deployment of these methods in real discovery programs, no pricing, licensing or downstream usage disclosure.
Field-level claims outrun measured performance on two of three tasks
The paper is itself framed as a reality check: the authors argue that most antibody design benchmarks are retrospective and that method claims proliferate in preprints and technical reports. Against blinded measurement, only affinity maturation performed broadly well; HCDR3 ranking was worse than random except for one model, out-of-library design frequently failed to beat standard selections, and successes did not transfer across tasks. That gap between general capability claims in the field and narrow demonstrated competence is positive and material. The gap is not larger because the paper's own conclusions are measured and it discloses its structural limits.
Organizing consortium reports and grades its own challenge
The paper discloses that the same consortium that ran the challenge is reporting the results and that blinding relied on organizer integrity rather than informatic controls, which is a structural incentive exposure for a benchmark whose purpose is to arbitrate competing method claims. Mitigating factors are the disclosure itself, the use of independent laboratories for assays, the release of underlying datasets, and results that are unflattering to much of the field rather than promotional. The sources disclose no funding, commercial ownership or participant-selection interests, so no inference is made about those.
Strong method transparency, single publisher and single round
Confidence is supported by the specificity of the reported design and readouts and by the authors' explicit statement of limitations. It is capped because the cluster contains one source, that source is the organizing consortium's own report, only one antigen and one round were evaluated, and the supplied excerpt does not break results down by named participant or model.
science
Flow control gets a shared benchmark, and a 38% friction cut nobody had to simulate first2 distinct publishers
build
Vivodyne is spending venture money on wet-lab throughput, not bigger models2 distinct publishers
science
Viral RNA snippets triple linear mRNA half-life in cells, IBS team reports in Cell1 distinct publisher
product
Vivodyne says the AI drug bottleneck is human tissue data, not model capability1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.