Science1 publisher3 min readPublished
AI antibody design clears one task in three in first blinded benchmark
AIntibody tested 511 antibodies from 29 organizations on three tasks. Ranking clones inside HCDR3 clusters was worse than random for every model but one.
The Scientist · Science desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- The AIntibody challenge tested 511 artificial intelligence-designed or predicted antibodies from 29 organizations, validated with diverse experimental assays and anchored to experimentally measured affinity and developability, in a prospective, blinded benchmark.
- AIntibody tested three tasks: in silico affinity maturation from phase 1 sequencing outputs, affinity ranking within heavy-chain complementarity-determining region 3 (HCDR3) clusters of a selection output, and CDR design of proteins not included in a selection output.
- AIntibody was a challenge inspired by the Critical Assessment of Structure Prediction (CASP), and represents a first step toward CASP-caliber benchmarking in antibody discovery.
- Several groups in the challenge produced developable antibodies with affinities below 100 pM.
- Affinity-matured antibodies were modeled effectively in the challenge.
Compiled by The ScientistSomething wrong?How this is made
Why it matters
A consortium has published what it describes as a prospective, blinded benchmark of computational antibody design, testing 511 AI-designed or predicted antibodies from 29 organizations against experimental affinity and developability readouts [1]. The result matters for procurement more than for research: performance tracked the task, not the vendor, and two of the three tasks exposed failure modes that no retrospective evaluation would have shown [5][6][7].
The challenge, called AIntibody and explicitly modeled on CASP, set three problems: in silico affinity maturation from phase 1 sequencing outputs, affinity ranking within heavy-chain CDR3 (HCDR3) clusters of a selection output, and CDR design of proteins that were not in a selection output [2][3]. Affinity-matured antibodies were modeled effectively [5]. Ranking high-affinity clones inside clustered HCDR3 datasets was, with the exception of a single model, worse than picking clones at random [6]. Out-of-library design was highly variable across most submissions, with many failing to outperform standard selection campaigns [7].
Several groups did produce developable antibodies with affinities below 100 pM [4]. According to the organizers, those successes were exceptions that did not transfer across tasks [8], which is the finding with the most direct commercial consequence. Only one of the three tasks yielded broadly reliable performance [14]. Across the field, 511 submitted antibodies from 29 organizations works out to roughly 18 per group [13], so this is a survey of many methods at shallow depth rather than a deep audit of a few.
The reason task specificity bites here is that antibody engineering is multiparametric. Beyond affinity and kinetics, developability covers expression, melting temperature, polyreactivity and aggregation, and the difficulty is improving one property without degrading others [9]. The authors put the asymmetry plainly: there are many solutions to improve a specific antibody, and many more ways to destroy it [10]. A model that ranks well inside a matured lineage is not thereby a model that can rank clones in a selection output.
The organizers are candid about the benchmark's own limits. This first iteration evaluated a single antigen, was organized by the same consortium now reporting the results, and relied on organizer integrity rather than informatic safeguards for blinding [11]. That is a weaker structure than CASP, which evaluated hundreds of diverse protein targets across multiple rounds under independent governance [12]. It is still stronger than the status quo, which the same authors characterize as performance claims resting on retrospective analyses without additional experiments, increasingly in preprints and technical reports [15].
The operating implication is to scope pilots by task with a defined baseline. For a ranking task, the baseline is random clone picking; for de novo CDR design, it is your existing selection campaign [6][7]. A platform that clears one bar tells you nothing about the other.
What to watch: whether future AIntibody rounds add antigens, independent governance and informatic blinding, which the organizers say they intend to address [16]; whether the sequencing datasets released with the study, which the authors offer for comparative benchmarking, get used by groups that did not compete [17]; and whether the named gaps in affinity prediction, library-inspired design and cross-task generalization close in a second round [18]. Related efforts, including the Adaptyv EGFR binder competition, which is not restricted to antibodies, provide a second reference point [19].