Science1 publisherNot yet confirmed elsewhere3 min readPublished
Calibrated metrics put AI virtual cell models ahead of simple baselines
Shift Bioscience reports in Nature Biotechnology that deep learning perturbation models often beat simple baselines once the metric can detect a biological signal. A model that looked no better than a baseline may just have been scored with a metric too blunt to tell.
The Scientist · Science desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Using the better-calibrated metrics, the team benchmarked nine deep learning models at predicting unseen single genes and unseen gene combinations.
- Judged that way, earlier models such as scGPT and GEARS cleared the uninformative baselines they had only seemed to match.
- Newer systems including PRESAGE and scLambda scored higher still, though their margins shifted with the dataset and the metric.
- The underlying metric audit ran 18 evaluation metrics across 14 genetic perturbation datasets.
Compiled by The ScientistSomething wrong?How this is made
Why it matters
- constraint A better metric cannot create signal that is not there; the gains over a simple additive rule showed up only where training data sampled a far larger share of gene combinations.
- capability If calibrated screens hold up in the lab, teams could triage candidate drug targets in silico before committing wet-lab time.
- precedent A benchmark for perturbation models now needs a prior check that its scoring metric can separate a known signal from noise.
A metric is calibrated, in this framework, when it can tell a real signal from noise. The test is plain. Build a positive control that stands in for a near-perfect prediction, build a negative control with no perturbation signal, and ask whether a given metric separates the two. That separation score is the dynamic range fraction, or DRF. [5] For the positive control the team used an interpolated duplicate: the average perturbation profile combined with an independent technical replicate, each gene weighted by how strongly the evidence says it was affected. [4] A metric that cannot pull the positive control apart from the negative one cannot tell you whether a model beat a baseline either.
That reframes the earlier disappointment, the one where sophisticated models kept failing to beat simple baselines. [2] Mean squared error and control-referenced Pearson correlation, two common choices, scored poorly on calibration, worst of all on datasets where the perturbations were weak. [7] The measures that held up were weighted and rank-based ones such as WMSE, weighted R-delta-squared and normalized inverse rank, which by design give more weight to the genes a perturbation actually moved. [8]
A calibration result tells you about the metric, not the biology. It shows a measure can tell signal from noise, but not whether a given model's predicted signal is real. The comparisons here run inside existing datasets, on perturbations held out from training. [9] Clearing a calibrated bar on one of them is a research result. Whether a screen then finds drug targets that hold up at the bench is a separate question.
Metric choice is only half of it. In Norman19, where the training data covered just 0.63 percent of possible gene combinations, a plain additive baseline stayed hard to beat. [12] In Wessels23, with 10.3 percent coverage, roughly sixteen times as much, several models beat that baseline. [13][19] The authors read that as deep learning outperforming simple additivity once it is trained on a broader slice of the combinatorial space. [13]
Brendan Swain, Shift's chief scientific officer and founder, framed the payoff for his own company. [14] "Our findings show that by using well-calibrated metrics and the right dataset, virtual cell models can generate biologically meaningful insights," he said. [15] He was direct about where it leads: "We are applying this framework directly in our target identification program, focusing on targets whose inhibition can support both rejuvenation and treatment of age-related disease, giving us a clearly defined route towards clinical development." [16] The Cambridge company says it will start with fibrosis. [17] The framework comes from a firm with a commercial interest in the models it rehabilitates. That is a reason to want it reproduced by groups that are not Shift.
What to watch
- Whether groups unconnected to Shift Bioscience reproduce the DRF calibration framework on their own datasets.
- Whether Shift's fibrosis target screens yield in silico hits that validate in vitro and in vivo.
- Whether PRESAGE and scLambda keep their edge as more datasets are re-scored on calibrated metrics.