Skip to content

Product1 publisher3 min readPublished

AI BioDesign promises a database of molecules evolution never made

The Allen Institute, the University of Washington and Fred Hutchinson are building models and databases of designed proteins. David Baker told WIRED the field usually sees that something works before it knows why.

The Product Desk · Product desk

Photograph accompanying AI BioDesign promises a database of molecules evolution never made
Photo: wired.com

What happened

  • AI BioDesign pairs machine learning with large-scale laboratory work to design and test molecules and biological functions that do not exist in nature but are physically and chemically possible.
  • The Allen Institute, a Seattle biomedical research nonprofit, leads the project with the University of Washington and the Fred Hutchinson Cancer Center.
  • Illustrations named in the project include drugs for cancer and neurodegenerative disease, and enzymes that break down plastics in the ocean.
  • David Baker, who shared the 2024 Nobel Prize in Chemistry for computational protein design, is the project's chief scientific officer.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • decision A lab choosing whether to plan around this is choosing a dependency with no published price, schedule or licence, so the only thing available to commit to is stated intent.
  • constraint Because the output is a data and model layer, a group that wants a molecule still funds its own synthesis, assays and validation; the project moves where the work starts and leaves the amount of it alone.
  • precedent Baker's works-first ordering pushes reviewers and regulators toward accepting experimental evidence in place of a known mechanism, which is a bar other design groups will cite.
  • exposure The environmental risk from a plastic-degrading ocean enzyme lands on whoever authorises release; the lab that screened the design in containment does not carry it.

A designed protein that works before anyone can say why is the normal first result in this work, according to Baker. "The first step is usually observing that something works, and we often only later understand why," he told WIRED, adding that machine learning "is really good at that first stage" [8][9]. For anyone who has to defend a result to a reviewer, that ordering matters: the evidence comes first, and the explanation sometimes follows.

The pitch is large. Baker said "We're working to build a world where a cure for a new disease is created in weeks, not decades," and described crops thriving in conditions that once killed them and molecular machines pulling critical minerals from waste [10][11]. The described output of AI BioDesign is narrower: databases, models and tools the group calls seeds for the medicines and technologies of the future [3]. Those are two different things to plan around: a therapy, and a set of starting points that somebody else has to test in a lab.

Take "decades" at ten years and "weeks" at four, and the stated goal implies a compression of about 130 times [16]. The interview does not include a baseline development timeline, a budget, a schedule or terms of access to the databases [15]. So the multiple is an ambition. There is nothing in it for a planner to hold. The shape of the deliverable is firmer: a database and a model are either there for a lab to query or they are not.

On risk, Baker said "the primary risks are not much different than those associated with any new biological technology," and listed unintended interactions with living systems, unexpected environmental effects and misuse [6]. His answer on control is procedural: screen computationally before a molecule is synthesised, then test in contained laboratory settings and increasingly realistic validation before any real-world application [7][18]. That sequence fits a cancer drug. An enzyme designed to break down plastic in the ocean has to leave containment to do the job it was designed for [4].

WIRED's write-up attributes to an unnamed leader in the field a comparison between these methods and industrialization, electrification and the digital revolution [13]. The claim underneath it is narrower and easier to test. Biology has spent its history studying what billions of years of evolution produced, and this project intends to work through what is physically and chemically possible instead [1][12].

For a team deciding whether to write this into a grant or a roadmap, two questions decide it: whether you can name the file you would download on day one and its format, and whether you can name the assay you would run on it in week two, in your own containment, with your own hands. If the honest answer to both is a paper at some future date, the project belongs on a reading list. AI BioDesign answers the first in principle, and has not published terms [3][15].

What to watch

  • Whether the group publishes licence and access terms for the databases and models, and whether outside labs pay.
  • Whether a designed enzyme intended for ocean plastic gets a named regulator for environmental release.
  • Whether a designed molecule reaches a clinical trial with no published mechanism, and how reviewers handle it.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories