Science1 publisherNot yet confirmed elsewhere2 min readPublished
Tangermeme gathers the after-training analyses of genomic deep learning models into one toolkit
Tangermeme's authors have built one toolkit for the work done with genomic deep learning models after training, such as scoring variants and finding motifs. They argue that this work is largely independent of a model's architecture, so one package can replace the bespoke code that labs ship with each model.
The Scientist · Science desk

What happened
- According to the authors, no optimized repository for post-training analysis worked across model types, so analysis code usually shipped bespoke with each released model.
- Existing packages either implement one or a few algorithms, as Captum does, or focus on training and fine-tuning, as Selene, kipoi, CREsted, EUGENe and gReLU do.
- Tangermeme leaves loading and defining models to the user, so it is tied to no particular repository and to no group's choices about how models are stored and distributed.
- Attribution methods such as DeepLIFT/SHAP, saturation mutagenesis or a custom operation can run on an edited sequence in place of a plain prediction.
Compiled by The ScientistSomething wrong?How this is made
Why it matters
- decision If variant scores really are largely indifferent to convolutions versus transformers, a lab's choice of architecture no longer dictates which analysis code it has to write or adopt.
- cost Each lab that releases a model currently writes and maintains its own interpretation code; a shared package moves that work into one implementation others can inspect and reuse.
- capability Because the package is indifferent to how a model is stored or distributed, the same marginalization or variant scan can be run on models from different groups and the outputs compared directly.
The case for one shared package rests on a claim in the paper's introduction. The authors write that downstream use of a genomic model is "usually agnostic to model architecture or training" [3]. Their example is variant scoring. They say it is largely unaffected by whether a model uses convolutions or transformers, or by its optimizer and learning rate [3]. They add that the operations inside models, and the ways models are trained, vary more and change faster than what researchers do with them afterward [7].
Most of the analyses are one experiment run with different edits. In silico marginalization drops a short sequence, such as a motif from a database, into a template and compares the model's predictions before and after [9]. Ablation does the opposite: it alters a region of the input that the model may be responding to [9]. Variant effect estimation makes one or a few substitutions, usually noncontiguous, and compares again. The authors use it to fine-map candidate drivers [10]. The paper's example of combining the pieces is to ablate the nucleotides around a known motif and study how context influences transcription factor binding [12].
The thing this doesn't tell you is whether the model has the biology right. Each of those operations compares a model's output with its own output on an altered input [9][10]. If a motif moves the prediction, the model relies on that motif. Whether a cell relies on it has to be tested in cells.
The paper calls tangermeme "highly optimized" [1]. All of its work happens at inference: making predictions and calculating feature attributions on models that are already trained [8]. Training cost does not come into it. For a lab planning a large variant screen, the cost it pays is inference. The excerpt does not include runtime benchmarks for the optimization claim, or the side-by-side comparison behind the architecture claim.
The authors describe predictive accuracy as "only the first step toward biological discovery" [14]. The models they have in mind predict transcription factor binding, chromatin accessibility, alternative splicing, miRNA binding and RNA degradation rates, some at single-cell or spatial resolution [2]. "The ability to train or adapt models to new settings is invaluable, but support for using these models after training is more limited," they wrote [13].
What to watch
- Independent runtime benchmarks of tangermeme against the bespoke analysis code shipped with existing models, run on the same inputs.
- A direct test of the architecture claim: one variant set scored by convolutional and transformer models through identical tangermeme pipelines.
- Whether new genomic model releases ship tangermeme-based analyses in place of their own interpretation code.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence50
- Adoption
- Insufficient
- Hype gap+15
- Incentives45
- Confidence55
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
The authors describe tangermeme as a "highly optimized toolkit for 'everything-but-the-model' when it comes to genomic deep learning" and show how it can distill learned cis-regulatory patterns from models into human-interpretable insights.
- [2]
Deep learning models can predict transcription factor binding, histone modification, chromatin accessibility and architecture, transcription, alternative splicing, miRNA binding and RNA degradation rates, including at single-cell or spatial resolution.
- [3]
The authors write that downstream usage is "usually agnostic to model architecture or training"; estimating variant effects is largely unaffected by whether a model uses convolutions or transformers, or by optimizer choice or learning rate.
- [4]
According to the authors, an optimized repository for post-training methods that generalizes across model types does not yet exist; bespoke analysis code is typically packaged with models on release.
- [5]
Captum implements one or a few related algorithms, while Selene, kipoi, CREsted, EUGENe and gReLU focus on training and fine-tuning, some including model zoos of pretrained models.
- [6]
Tangermeme leaves the details of loading and defining models to the user so as not to be locked in to any particular repository or to others' design choices about how models are stored and distributed.
- [7]
The authors say the computational operations within models and their training strategies are more diverse and rapidly evolving than their downstream use.
- [8]
Tangermeme has modular implementations of sequence manipulations and of operations involving models, such as efficiently making predictions or calculating feature attributions, and these can be stacked to create analyses.
- [9]
In silico marginalizations compare predictions before and after substituting a usually short sequence into a template and can quickly identify relevant motifs from a database; ablations compare predictions before and after altering a region of the input the model may be responding to.
- [10]
Variant effect estimations compare predictions before and after one or a small number of usually noncontiguous substitutions and can fine-map candidates to identify drivers.
- [11]
Any model operation, including DeepLIFT/SHAP, in silico saturation mutagenesis or a custom operation, can be performed after sequence manipulation in tangermeme.
- [12]
The influence of context on transcription factor binding can be studied by ablating the nucleotides surrounding a known motif.
- [13]
"The ability to train or adapt models to new settings is invaluable, but support for using these models after training is more limited."
- [14]
The authors write that although most attention has been paid to maximizing predictive performance, this is "only the first step toward biological discovery".
Sources
1 independent publisher whose own reporting we read for this story.
- nature.comTangermeme: a toolkit for understanding <i>cis-</i>regulatory logic using deep learning models
1 article · October 7, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.
Topics
- Variant effect predictionFollow
- Model InterpretabilityFollow
- Cis-regulatory genomicsFollow
- Genomic deep learningFollow