Skip to content

Science2 publishers2 min readPublished Updated

Berkeley's genome model buys accuracy with alignments instead of parameters

GPN-Star folds whole-genome alignments and a species tree into its architecture, which is how it trains in hours where Evo 2 needed 2,000 processors and months. The cost of building those alignments sits outside that comparison.

The Scientist · Science desk

Photograph accompanying Berkeley's genome model buys accuracy with alignments instead of parameters
Photo: nature.com

What happened

  • UC Berkeley's GPN-Star, published in Nature, is a genomic language model whose architecture takes whole-genome alignments and species trees as explicit inputs rather than learning from unaligned genomes.
  • Trained on alignments spanning vertebrate, mammal and primate timescales, it reports state-of-the-art variant effect prediction across both coding and non-coding regions of the human genome.
  • The paper also reports substantial gains over previous methods in prioritizing pathogenic and fine-mapped GWAS variants, in complex trait heritability enrichment and in rare variant association power.
  • Training took days or even hours on a handful of processors, against the 2,000 NVIDIA processors and months of runtime that the larger Evo 2 model required.
  • Versions were trained for five model organisms as well: mouse, chicken, Drosophila melanogaster, Caenorhabditis elegans and Arabidopsis thaliana.

Compiled by The ScientistSomething wrong?How this is made

Why it matters

  • cost The cheap training run sits on top of an expensive input: high-quality whole-genome alignments were expanded by large-scale consortia, so the compute saving is partly a transfer of cost to whoever builds and maintains the alignments.
  • constraint Because the model's signal is borrowed from comparative genomics data, its usefulness travels only as far as good alignments already do, which is a real limit for any organism that lacks them.
  • decision A group choosing a tool is not weighing equivalents: Evo 2 can generate whole genomes, while GPN-Star scores existing variants, and only the scoring job gets cheaper here.
  • capability With genome-wide predictions released, a lab that can afford ten assays instead of ten thousand now has a published ranking to spend them against.

The denominator first: of the human genome's 3 billion base pairs, 1% to 2% codes for proteins [17], which leaves roughly 2.94 to 2.97 billion bases of evolutionary holdover and regulatory sequence [18]. Working out which of those bases matter is what a conservation signal is for. A whole-genome alignment relates hundreds of species to one anchor genome and marks, site by site, where sequence has been conserved and where it has changed [7]. A model trained on unaligned genomes has to recover that correspondence itself; GPN-Star receives it as input, and that is where the training savings come from [8].

The Nature paper is direct about the problem this addresses. Genomic language models built on standard language-modelling frameworks still lose to much simpler classical phylogenetic models on some variant interpretation tasks even at massive model sizes, and the shortfall is worst in complex eukaryotic genomes and in distal regulatory elements such as enhancers [9]. PhastCons and PhyloP, the parametric models fitted on alignments, have been staples of human variant interpretation for years [10]. Proteins met this argument first: alignment-based models such as AlphaFold, MSA Transformer and EVE worked, and interest in alignments revived once scaling single-sequence protein language models showed diminishing returns [11].

Evo 2 ingested more than 100,000 species using 2,000 NVIDIA processors over months [5]. The alignments GPN-Star learns from span hundreds of species [7], so even reading "hundreds" at its upper bound, the alignment route looks at fewer genomes by a factor of at least a hundred [19]. The efficiency claim is about how much raw sequence ever enters the model, not only about parameter count.

Neither the human-genome benchmark result nor the prioritization gain comes with a number in the abstract; the wording is qualitative [21]. Yun Song, the study's senior author [22], places the model upstream of the bench, saying that nobody can experimentally test every variant in the genome and that the predictions should help prioritize the experiments with the greatest potential impact on human health [16]. For a conservation-derived score, that is the defensible position: it ranks candidates, and something else confirms them.

One result deserves more attention than the ranking. Modelling recent evolution and modelling deep evolution help different tasks, and the paper reports that advantage as task-dependent [13]. The vertebrate, mammal and primate versions are therefore three different instruments rather than three rungs of a ladder, and choosing among them falls to whoever runs the model on a particular question.

What to watch

  • Independent benchmarks that put a number on GPN-Star's margin over PhyloP, PhastCons and Evo 2, especially in enhancers.
  • Experimental follow-up on the published genome-wide predictions: which prioritized variants survive assay.
  • Whether the framework transfers to species without high-quality whole-genome alignments already built.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories