Skip to content

Written by AI.How we work

Invest3 publishersIndependently confirmed3 min readPublished

Meta, DeepMind and Isomorphic pay $300 million for first use of Biohub's open cell data

Meta, Google DeepMind and Isomorphic Labs put $300 million into Biohub's $1.8 billion Virtual Biology Initiative and get the data before the public does. Startups that sell proprietary cell data will now be priced against a free dataset that their best-funded rivals see first.

The Investor · Invest desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Photograph accompanying Meta, DeepMind and Isomorphic pay $300 million for first use of Biohub's open cell data
Photo: yahoo.com

What happened

  • Biohub, the philanthropic venture of Mark Zuckerberg and Priscilla Chan, committed the first $500 million to the project in April.
  • Over five years, the Department of Energy is putting more than $500 million toward lab measurement, modeling and computing work.
  • Datasets paid for with more than $500 million of earlier federal funding will be coordinated by the National Institutes of Health. Biohub will standardize them for AI training.
  • A January dataset made with Tahoe Therapeutics and Arc Institute covers more than 120 million single cells and 225,000 perturbation interactions.

Compiled by The InvestorSomething wrong?How this is made

Why it matters

  • exposure Startups whose pitch rests on proprietary cell data must now compete with a public set that Isomorphic Labs, itself a drug-discovery company, gets to work on first.
  • decision Each drugmaker Biohub approaches next has to decide whether embargoed early access is worth paying for or whether to wait for the free public release.
  • contradiction Cryptobriefing reports that all data and models will be openly accessible, while Rives describes embargo periods for commercial funders before the data becomes public.

The four commitments behind Biohub's headline add up to $1.8 billion [18], give or take the "more than" attached to both federal figures [3][4]. One of them, the NIH's, is federal money already spent building repositories that exist today [4]. Take it out and new commitments come to roughly $1.3 billion [22]. The Department of Energy's portion works out to more than $100 million a year [20]. Meta, Google DeepMind and Isomorphic Labs, with $300 million between them [2], supply about 17% of the total [19].

That 17% buys time [19]. "With commercial funders we have embargo periods where there's a period of time where the groups can work on the data, and then it becomes available as a public scientific resource," Alex Rives, Biohub's head of science, said [13]. For $300 million the three companies get a lead on data that will eventually be public, and they split the cost of producing it with Biohub and two federal agencies [2][3][4]. Biohub did not say how long the embargo runs or how the $300 million divides among the three, and neither report cites a startup financing that would show the effect on venture money [13].

"We have always held this as a community asset, not just for one group, so that it can build upon itself over time," Priscilla Chan said [9]. With Meta also writing a check, Cryptobriefing noted, the Zuckerberg name appears on both sides of the funding [11]. For biotech investors the funder that matters is Isomorphic Labs, which the CNA report calls a drug discovery startup [12]. A company in the business of finding drugs gets to work on the shared data before its competitors do [13][12].

Scale is still the binding constraint. Current cell datasets run to hundreds of millions of cells, and Rives said an accurate predictive model will need billions and eventually trillions [6]. Reaching one billion would take more than eight times the largest set the initiative has produced so far [21]. Rives expects a first dataset in about a year and accurate predictive models within five [7].

For startups that raise money on proprietary cell data, the embargo length decides most of the outcome [13]. If it is long, the three funders have bought a private lead for $300 million [2], and startups face better-funded rivals who see a larger set first. If it is short, the public baseline rises each year and a private perturbation screen is worth less as a fundraising asset. And if the data reaches billions of cells late, private sets keep their value for longer [6]. I think the short-embargo case is the most likely one for startups whose data resembles Biohub's, which comes from spatial transcriptomics and screens of how cells respond to changes in environment [8]. The federally funded work running alongside has no embargo [14].

The counter-thesis is that cell-level data was never where these startups kept their value, and that companies holding other kinds of data keep their pricing. Other labs are moving into biology data too: Anthropic has added a wet lab, and the OpenAI Foundation has started a grant program of more than $125 million for biological and medical datasets for AI research [17]. If the commercial embargo turns out to run for years, I am wrong, and the advantage goes to whoever can pay for early access [13].

What to watch

  • The length of the commercial embargo, if Biohub or a funder discloses it: months of exclusive use and years of it are different purchases for $300 million.
  • The size of the first dataset when it ships, measured against the billions of cells Rives says an accurate model needs.
  • Any split of the $300 million among Meta, Google DeepMind and Isomorphic Labs, which would show how much the one drug-discovery company paid for its head start.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories