Invest2 publishers3 min readPublished Updated
DeepMind makes Isomorphic Labs license the genome atlas academics can browse free
Academics get all 9 billion single-base predictions through a browser at no cost, while commercial users wait on a Google Cloud licence whose terms DeepMind has not published, including for its own sister company.
The Investor · Invest desk

What happened
- Google DeepMind said Tuesday it has predicted the biological consequences of all 9 billion possible single-letter changes to human DNA and is releasing the resulting database free to academic researchers.
- Pushmeet Kohli, DeepMind's vice president for research, told reporters it is the first time any researcher in the world can reach a comprehensive map of human genetic variation by simply opening a browser.
- Non-commercial use starts today through a website DeepMind has set up, while commercial access is promised soon through a licensing arrangement on Google Cloud.
- Kohli said sister company Isomorphic Labs will have access to the Atlas but will also require a commercial licence, and he did not specify what those terms would be.
- The release includes an AlphaGenome Variant Impact score that merges gene-regulation predictions with the earlier AlphaMissense protein model, where 10 marks the top decile and 30 the top one in a thousand.
Compiled by The InvestorSomething wrong?How this is made
Why it matters
- constraint About 9 million variants sit in the top one-in-a-thousand AVI band, far more candidates than any wet lab can assay, so the binding limit on discovery becomes experimental throughput rather than model compute.
- decision Anyone currently paying to score variants in house faces a build-or-license choice they cannot yet cost, because the Google Cloud terms do not exist in public.
- precedent A licence written between two Alphabet subsidiaries sets scientific reference data up as a metered Cloud product line, which is the template the next precomputed dataset will be distributed under.
- exposure Academic groups wiring Atlas lookups into their pipelines are betting on continued free non-commercial access under a licensing arrangement that is still being written.
Three billion reference bases, each with three alternative letters, is how you arrive at 9 billion [9][1], and at an average of about 27,000 predictions per variant across hundreds of human and mouse cell and tissue types [10], the catalogue carries roughly 243 trillion individual predictions [2]. DeepMind paid that compute bill once and has not said what it came to. The marginal cost of the next lookup, to the researcher doing the looking, is a page load [6]. The comparison DeepMind offers is the old workflow of running a model one variant at a time or testing in a laboratory [3]; the comparison a large buyer will actually run is against its own hardware bill for scoring variants in house, and that is the number the unpublished licence terms will be measured against [6].
Exhaustiveness stops at substitutions. The same release scores more than 100 million insertions and deletions, but those are variants observed in population databases including the UK Biobank and the NIH's All of Us [11], which is about 1.1% of the substitution count [4] and a different sort of object: what has been seen in people, rather than what is chemically possible. Chase an indel and you are back to asking whether anyone has recorded it.
The internal licence matters most here. If Alphabet's drug-discovery arm has to take a commercial licence to use a dataset produced by its sibling [7], Atlas is being run as a priced product line rather than a shared corporate asset, and the free academic tier is what establishes it as the reference everyone benchmarks against before the paid tier has a number attached. That is the thesis, and the counter-thesis sits inside the same missing figure: a nominal licence written for accounting hygiene and a seven-figure one imply opposite answers about where value accrues, and nothing announced distinguishes them [7]. How it plays out depends on the pricing. If the licence is set to displace in-house variant pipelines, Google books the difference. If it functions instead as a funnel, the Cloud compute around the data ends up earning more than the data itself. And if it is priced above what pharma already pays to run its own models, Atlas settles into being an academic public good with a toll booth nobody walks through.
Other than published terms, the thing that would settle it is validation. The account of how Atlas was built is a bioRxiv preprint [8], and the claim being made for the exercise is Kohli's, that the Human Genome Project bought the book in 2003 without learning to read it [5]. A reading aid gets judged on whether the reading changed, which shows up as confirmed variant-to-disease links, not as a prediction count.
What to watch
- The Google Cloud commercial licence terms, and whether the rate Isomorphic Labs pays is ever disclosed.
- Whether the bioRxiv preprint survives peer review with validation strong enough for clinical variant interpretation.
- Whether free non-commercial access stays uncapped once a paid tier exists alongside it.