Skip to content

Product3 publishers3 min readPublished

DeepMind turns variant scoring into a lookup by precomputing 9 billion mutations

AlphaGenome Atlas answers every possible one-letter question in the human genome in advance, roughly a petabyte of predictions. Noncommercial access opened Tuesday, commercial use waits on Google Cloud, and the launch write-up carries no accuracy number.

The Product Desk · Product desk

What happened

  • Google DeepMind published AlphaGenome Atlas on Tuesday, a precomputed predictive map covering every possible single DNA letter change in the human genome, roughly 9 billion substitutions in all.
  • Alongside it Google released a Variant Impact Score, or AVI, which the company says lets researchers rank variants and read their molecular effects in the same step.
  • Researchers reach the predictions three ways: a web portal, a skill inside Google's agentic development platform Antigravity, and the existing AlphaGenome interface.
  • Access splits by use: noncommercial researchers get Atlas through the website from launch day, while commercial use is promised on Google Cloud only "soon".
  • Atlas extends beyond AlphaMissense, which predicted protein-altering mutations, into the noncoding stretches of the genome that regulate how and when genes switch on.

Compiled by The Product DeskSomething wrong?How this is made

Why it matters

  • decision A genomics team's next call is no longer how much compute to buy to score candidate variants, but whether it will act on a table someone else computed and it cannot rerun.
  • constraint Anyone building variant interpretation into a product cannot put Atlas in the pipeline yet, and has no date to plan the procurement around.
  • exposure A lab that drops a candidate because AVI ranked it low is carrying a judgment made by a model whose error rate was not published with the launch.
  • capability Regulatory-region candidates that previous DeepMind tools could not speak to now get a first-pass answer before anyone books bench time.

The person this lands on has a list of letter changes off a sequencing run and no reliable way to say which of them matter, which is exactly the problem DeepMind names at the centre of genomics [12]. Until this week that problem was shaped like compute: obtain the model, find the hardware, score your candidates. Ziga Avsec, DeepMind's genomics lead, was straightforward about why the catalogue lagged the model in the press briefing: "Basically it took us some time to really precompute and also analyze this many variants because the space is so big" [9]. AlphaGenome itself came out last year [10].

The arithmetic behind the headline number is simple enough to check. Each position in the genome has three other letters it could become, so roughly 3 billion letter pairs produce roughly 9 billion possible substitutions [2][3][1]. Google puts the resulting dataset at about a petabyte [7], which averages out near 110 kilobytes of prediction per variant [2]. Nobody pulls that down to a cluster, since the whole point is that you query it instead.

Google describes this as transforming our understanding of biology and paving the way for new treatments [14]. What has actually shipped is a very large table with a sort column: per variant, an estimate of molecular consequence such as how much of a protein gets made [4], plus the Variant Impact Score for ranking [6]. That table was built by a model trained on public human and mouse genome databases [11], and The Verge's account of the launch and briefing includes no accuracy or validation benchmark [15]. The catalogue answers what a change might do at the molecular level, but not whether the change matters in your patient or your programme, which DeepMind still frames as the open question [12].

So the honest answer to who can use this on Monday is a researcher in an academic or nonprofit lab with a browser [8]. Everyone whose interpretation step sits inside a commercial pipeline is waiting on a licence with no date attached [8].

For those who can query it, a two-by-two beats a hit list. One axis is the Atlas call, high impact or low. The other is whether you hold local evidence on that variant. High impact with no local evidence is where a wet-lab slot earns its keep. High impact where your own data disagrees is the cell worth writing up, because it tells you something about the model rather than the variant. Low impact with local agreement is cheap confirmation. Low impact with no local evidence is where the quiet decision happens: you will drop a candidate on a prediction alone, so record that you did, because "we never looked" and "the score said no" are different positions to defend later.

Google paid the precompute bill once. The cost of being wrong about any single variant still lands on the lab that cites it.

What to watch

  • The licence terms and an actual date for commercial Atlas access on Google Cloud, currently dated only as "soon".
  • Whether DeepMind publishes benchmark accuracy for Atlas calls against experimental data, and separately for noncoding regions.
  • Whether any clinical or curation workflow starts accepting AVI scores as evidence for deprioritising a variant.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories