Skip to content

Science1 publisher3 min readPublished

DeepMind precomputed all 9 billion single-letter genome changes into a free petabyte atlas

The AlphaGenome Atlas replaces a compute-heavy programming interface with a browser lookup for every possible single-base change in human DNA. The evidence that its impact score flags disease-linked mutations sits in two preprints.

The Scientist · Science desk

Photograph accompanying DeepMind precomputed all 9 billion single-letter genome changes into a free petabyte atlas
Photo: nature.com

What happened

  • Google DeepMind announced AlphaGenome Atlas on Sept. 8, a free database of 9 billion possible changes to the human genetic code with estimates of their effects on different tissues and cellular processes.
  • DeepMind computed every possible base change in advance, producing 1 petabyte of output served through a web portal instead of the programming interface researchers previously had to run themselves.
  • A Sept. 11 preprint from Katie Pollard's team at the Gladstone Institute found AlphaGenome adept at identifying causal mutations while also making one error persistently.

Compiled by The ScientistSomething wrong?How this is made

Why it matters

  • capability A geneticist with no bioinformatics training and no cluster time can now retrieve a predicted effect for any single-letter change, so the group acting on these scores is larger than the group able to check how they were produced.
  • decision Labs writing variant prioritisation rules have to settle how far a precomputed score moves a candidate up the bench queue while the supporting evidence is two preprints.
  • constraint One summary number per variant limits what a triage pipeline can condition on, and Lappalainen said that simplicity can make the score harder to interpret in certain cases.
  • contradiction The strongest positive result comes from the developer's own paper while the independent one reports a recurring error, so how much confidence a lab assigns depends on which preprint it weights.

Each of the roughly 3 billion positions in the human genome can hold one of three other letters, so 9 billion is the complete set of single-base substitutions [1][2][1]. All of them are in the atlas. DeepMind ran the model at every position and stored the output, and the total comes to 1 petabyte, which works out at about 110 kilobytes of predicted consequence per variant [5][2].

Most of that lies outside genes. About 98% of human DNA is noncoding, which puts roughly 8.8 billion atlas entries in sequence that does not itself encode protein [11][3]. Geneticists once dismissed that fraction as junk and now treat it as the instructions governing when coding genes are read, with some of those instructions spread widely through the code [16]. AlphaGenome scores a change at one base pair for its influence on up to 1 million surrounding base pairs, a wider spread than earlier models managed [12]. Google DeepMind says the model also beats previous ones on the resolution of those predictions [10].

Before Atlas, using AlphaGenome meant working through its programming interface with enough bioinformatics skill to do so, and each variant calculation taxed academic computing resources [4]. "You really do not need to be an expert in these methods to be able to go there and look something up in a browser," said Tuuli Lappalainen, a genomics professor at KTH Royal Institute of Technology in Stockholm whose lab has spent the past year using AlphaGenome to study genome variation [6][9]. Greg Findlay, a group leader at the Francis Crick Institute in London who is not involved in the project, told Live Science: "It looks like a great resource" [7].

Two preprints carry the validation. DeepMind's accompanying paper showed the AlphaGenome Variant Impact score separating disease-linked mutations from harmless ones in a clinical dataset [13], a result that holds at the level of groups. The score distributions differ between the two labelled sets. A lab working from a shortlist gets no probability that a particular variant on it is the causal one. On Sept. 11, a team led by Katie Pollard, director of the Gladstone Institute of Data Science and Biotechnology and a professor at the University of California, San Francisco, posted a preprint that found AlphaGenome adept at identifying causal mutations while also making one error persistently; Live Science did not say what that error was [15].

Lappalainen said the score's simplicity might make it harder to interpret in certain cases, and that it would be useful for researchers who want to investigate a list of gene variants [14]. She also said the underlying technology is not new [9]. Asked to place it, she told Live Science: "It's not this holy grail." [8]

What to watch

  • Whether the two preprints survive peer review, and whether the persistent error Pollard's group reported appears in the published version.
  • Whether DeepMind publishes per-variant calibration for the Variant Impact score instead of group separation on a clinical dataset.
  • Whether independent labs report how the score performs specifically on noncoding variants, where most of the 9 billion entries sit.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories