Skip to content

Invest1 publisher2 min readPublished

Google DeepMind put 111 kilobytes of prediction behind each of 9 billion DNA variants

AlphaGenome Atlas scores every possible single-letter change in the human genome and hands the result to anyone with a browser, which moves the scarce input from computing a prediction to confirming one. Price and drug-developer partnerships go unmentioned in the release.

The Investor · Invest desk

Illustration accompanying Google DeepMind put 111 kilobytes of prediction behind each of 9 billion DNA variants

What happened

  • Google DeepMind released AlphaGenome Atlas, a one-petabyte dataset of molecular-effect predictions for more than 9 billion possible single-letter DNA changes, developed over several years.
  • Scientists reach the predictions through a web browser without writing any code, which is how the resource is intended to be used at scale.
  • Researchers at the Stowers Institute, Broad Institute, the University of Exeter, Memorial Sloan Kettering and Stanford contributed scientific input and tested applications alongside Google DeepMind.
  • The underlying research is out as a preprint rather than a peer-reviewed paper.

Compiled by The InvestorSomething wrong?How this is made

Why it matters

  • constraint By the release's own account, laboratory testing of every possible change is impractical, so the binding resource is validation capacity, which nine billion predictions cannot buy on their own.
  • decision Labs and biotechs currently funding in-house genome-wide variant scoring have to decide whether that spend still buys anything a browser query does not, and redirect it if not.
  • exposure Anyone whose product is a proprietary variant-effect score is now priced against a public resource whose licence and pricing terms are not stated, an exposure defined by an undisclosed term sheet.
  • precedent Academic institutions supplying expertise and feedback to a model owned by one company, with no disclosed funding or IP arrangement, sets the template for how the next such atlas gets sourced.

A petabyte spread across nine billion variants comes to roughly 111 kilobytes of prediction per single-letter change [15], which describes a dense per-variant record rather than a ranked shortlist, and the coverage is combinatorial rather than curated, since three billion letters with three alternative bases apiece is where the nine billion figure comes from [16][3].

The release is straightforward about why no laboratory has done this: testing the effect of each change would be practically impossible [4]. So the Atlas converts a search problem into a triage problem, and triage runs on wet-lab throughput and cohort access, neither of which a browser query supplies [7]. Google DeepMind's own framing is that until now no single resource both ranked variants genome-wide and named the biological processes they are predicted to disrupt [5], a claim about ranking; confirmation is a separate matter.

Asking a question here costs a browser tab and no code [7]; pricing, licence terms, and who funded what are absent from the announcement [17]. The nearest thing to a commercial channel in the document is a job title, since Pushmeet Kohli is described as VP Science at Google DeepMind and Chief Scientist at Google Cloud [20].

Five outside institutions are named, all of them research institutes or universities [6][19], and the division of labour runs one way in the description: Google DeepMind built the technology, the academic partners supplied biological expertise and feedback that guided it [8]. The document lists academic institutions rather than drug developers, and it does not claim any variant as validated; the underlying work is a preprint [9], so the conclusion that defensible value has already relocated to therapeutics is running ahead of the evidence in this document.

A more interesting version of the counter-thesis is worth holding onto: a predictor this broad raises the value of the functional-genomics measurements used to train and calibrate it, in which case data generation gets more valuable rather than less, and the party best placed to commission that data is the one that owns the model [8]. The reading the document supports is narrower than the launch invites, which is a browser-accessible triage layer, built and controlled by a single party [8], useful in strict proportion to its calibration, and calibration is what a preprint has not yet been shown to have [9]. What would settle it is a hit rate, meaning Atlas-prioritised variants tested against laboratory outcomes on held-out clinical cases. Nine billion predictions [2] remain nine billion hypotheses because the release itself explains why they cannot all be checked [4].

What to watch

  • A refereed version of the preprint carrying hit rates for Atlas-prioritised variants against laboratory outcomes.
  • Any change in access or licence terms, especially commercial use routed through Google Cloud.
  • The first drug developer to name the Atlas in a target-selection disclosure.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories