Skip to content

Leadership1 publisher3 min readPublished

Google precomputes 9 billion variant predictions into a one-petabyte lookup

AlphaGenome Atlas turns per-variant inference into a table lookup and adds a ranking score, which changes what a genomics budget buys this quarter well before it changes what a clinic can act on.

The Board Room · Leadership desk

Illustration accompanying Google precomputes 9 billion variant predictions into a one-petabyte lookup

What happened

  • Google DeepMind released AlphaGenome Atlas on Sept. 8, 2026, a searchable store of predictions for 9 billion possible single-letter changes in human DNA, held in a dataset of about one petabyte.
  • The accompanying AVI score folds AlphaGenome predictions together with AlphaMissense, evolutionary conservation and protein-coding features into a single rank for each variant.
  • An independent CRISPRi benchmark from Cold Spring Harbor Laboratory put AlphaGenome top on both tests of distal gene-control disruption, though its predictions understated measured effect sizes.

Compiled by The Board RoomSomething wrong?How this is made

Why it matters

  • decision The budget line that shrinks is inference per question; the one that does not is the bench work and analyst judgement that turn a rank into an answer, and the lab pays for that itself.
  • constraint With no clinical validation or approval behind the model, a rank can enter a diagnostic chain of evidence but cannot close it, so confirmatory assays stay on the cost sheet.
  • exposure The most persuasive case study comes from coauthors of the Atlas manuscript, so a leader weighing adoption is weighing collaborator evidence and should price it accordingly.
  • capability Groups with no inference budget can now triage the full substitution space before deciding which follow-up to fund, a step that used to be a compute purchase.

About a petabyte spread across 9 billion variants works out to roughly 110 kilobytes of stored data per variant [18], which matches a resource in which each variant carries contributing features a researcher can open and inspect before choosing an experiment [6]. What Google has built is a table, and the questions the table answers were fixed at the moment it was computed.

Precomputation buys breadth at the price of currency. The gain is that a lookup replaces a run for every possible single-letter substitution in human DNA [1][2]. The cost is that the table has releases, and any result inherits whichever release was reachable when the work was done. The Atlas team's own biobank analysis used slightly outdated Atlas releases because the UK Biobank computing platform was temporarily unavailable [12]. That is an operational footnote with a durable implication: version control over the reference table becomes part of the audit trail for anything derived from it.

The quantified gain is a prioritisation gain, and the manuscript says so. Applied to protein measurements from 54,189 UK Biobank participants after quality control, Atlas features produced a 22% increase in discoveries of statistically independent groups of rare non-coding variants associated with protein levels [10], and the source is explicit that this was not a gain in diagnostic accuracy [11]. For a research budget those are two different purchases: more hypotheses worth testing per unit of analyst time, or more confidence in an answer, and Atlas delivers the former, not the latter.

The strongest evidence for the DNM1 case comes from people with their names on the paper: Laura Covill, Anne O'Donnell-Luria and their Broad Institute colleagues are coauthors of the Atlas manuscript, which makes their work collaborator research rather than unaffiliated replication [9]. That evidence supports a workflow claim rather than a clinical one. AVI ranked a non-coding change in DNM1 first in an unresolved epileptic encephalopathy case and predicted that the change created a false splice site, adding an abnormal extension to a brain-specific version of the protein [7]; laboratory screens then reproduced the abnormal splicing and turned up nearby variants with similar effects [8]. A rank pointed at a mechanism someone could test, and the test worked.

The independent evidence is narrower than the headline number. Alan Murphy and Peter K. Koo of Cold Spring Harbor Laboratory found that AlphaGenome recorded the highest correlations on both of their CRISPRi tests of how disabling distant gene-control regions alters gene activity, in a near tie with Borzoi on the Fulco test [13], while its predictions understated the size of measured effects, with the gap widening for more distant regions [14]. Their benchmark ran in K562 cells that are well represented in model training, and it tested the underlying AlphaGenome model rather than the Atlas or AVI [15]. Directional ranking is what the outside record supports; effect magnitude and other cell types are not.

For decisions taken this quarter, the question is whether a variant triage step still needs its own inference budget, and for human single-nucleotide variants that answer has moved toward no [1][2]. For the decade, the question is whether predicted molecular effects can carry weight inside a diagnosis, and the record does not answer it: Atlas and AVI predict molecular effects that can support a chain of evidence without establishing a diagnosis on their own [16], and AlphaGenome has not been validated or approved for clinical use [17]. A programme that plans for the Atlas to order its experiments is planning on what exists; one that plans for it to close cases is planning on a level of validation the record does not yet contain.

What to watch

  • An unaffiliated group replicating the DNM1 splicing result without Atlas team coauthors on the paper.
  • A benchmark that tests AVI itself, and in cell types not well represented in AlphaGenome's training data.
  • Whether Google publishes release versioning that supersedes the Atlas releases used in the UK Biobank analysis.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories