Skip to content

Science1 publisher2 min readPublished

AbbVie, argenx, Lundbeck and Takeda send antibody sequences to Ginkgo for a 10,000-antibody training set

AbbVie, argenx, Lundbeck and Takeda are founding a Ginkgo and Apheris consortium to build a standardized developability dataset of 10,000 antibodies. Ginkgo will oversee production and wet-lab testing of the whole set, with members due to see the first data in early 2027.

The Scientist · Science desk

Illustration accompanying AbbVie, argenx, Lundbeck and Takeda send antibody sequences to Ginkgo for a 10,000-antibody training set

What happened

  • Each founding member contributes proprietary antibody sequences, and publicly available sequences fill whatever capacity the members leave.
  • Ginkgo is training a foundation developability model inside Apheris' secure environment, and Apheris delivers it into each member's systems for fine-tuning on private data.
  • Charlotte Deane of the University of Oxford and Peter Tessier of the University of Michigan will provide independent scientific oversight.
  • The consortium remains open to further pharmaceutical and biotech companies beyond the four founders.

Compiled by The ScientistSomething wrong?How this is made

Why it matters

  • exposure Sequences are shielded from rival members, but Ginkgo receives them in order to produce the antibodies, so each company's confidentiality depends on its terms with Ginkgo.
  • decision Drugmakers outside the founding four have to weigh handing sequences to Ginkgo against gaining a model and dataset that their competitors also use.
  • precedent If the standardized set does yield better models, the planned move to more complex antibody formats would carry the same shared-measurement approach into new drug classes.

Ginkgo Datapoints also designs the sequence selection, so which molecules enter the set is decided before anything is measured [8]. Athena Hadjixenofontos, who heads AI in biotherapeutics and genetic medicine at AbbVie, said the consortium is "creating datasets that are designed for machine learning, addressing limitations associated with convenience datasets" [5]. I take convenience data to mean measurements that already existed, collected in different labs for other purposes. When one organisation picks the molecules, oversees their production and runs the same high-throughput characterization on every one [8], a difference in a readout is much harder to blame on whoever took the measurement.

The partners claim the result will be the field's largest standardized antibody developability dataset [2]. The target is 10,000 antibodies [4]. The announcement does not say how many will be proprietary or which assays make up the core endpoints, and it does not quantify how far models trained on existing data fall short. The proprietary share matters because the case for pooling rests on it: if public sequences make up most of the set, any improvement would owe more to uniform measurement than to pooled pipelines. "Pooling standardized developability data across the industry can create stronger predictive models than any one company could build alone," said Yves Fomekong Nanfack, head of AI/ML research at Takeda [15].

The consortium's stated aim is to flag manufacturability and developability risks earlier [1]. The partners define developability as the biophysical properties that decide whether a candidate can be manufactured, formulated and advanced into a clinical product [7]. Members will benchmark their models against Ginkgo's wet-lab readouts [8][10]. The thing this set cannot tell a member on its own is whether a model that predicts those readouts also catches the candidates that later fail in that company's manufacturing or formulation work. Hadjixenofontos said the federated setup lets participants "contribute data while keeping proprietary sequences private" [6], so a company can fine-tune the shared model on its internal records [9] without showing them to rivals. Allan Jensen, vice president of biotherapeutic discovery at Lundbeck, said that in complex therapeutic areas such as CNS, "the ability to select well behaved candidates with superior developability properties is essential" [14].

What to watch

  • Any published comparison, for instance from advisers Deane and Tessier, of the foundation model's predictions against members' own manufacturing and formulation outcomes.
  • New pharma or biotech members signing on before the first data release, since each would add proprietary sequences to the set.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories