Science1 publisher2 min readPublished
Common perturbation-model metrics fail a positive control built from real cells
Researchers testing 18 metrics on 14 datasets found MSE and Pearson delta often miscalibrated, too insensitive to register real skill in perturbation models. That weakens recent verdicts that the models do no better than an average guess.
The Scientist · Science desk

What happened
- Several independent benchmarks had found that a mean baseline, the average perturbed expression profile in the training set, often matched or beat state-of-the-art models on MAE, MSE and Pearson delta.
- A positive control that predicts half of each perturbation's cells from the other half lost to that mean baseline on MSE in about 95% of perturbations in the Replogle22 K562 genome-wide dataset.
- That genome-wide dataset averaged 3.25 differentially expressed genes per perturbation, against 102.09 in Norman19, and the authors traced the control's failure to that dilution.
- The authors' replacement control weights each gene between the half-split and the mean according to its differential-expression P value, and it scored consistently positive.
Compiled by The ScientistSomething wrong?How this is made
Why it matters
- decision Groups weighing whether to invest in in-silico screens should not treat a model's loss to the mean baseline on MSE or Pearson delta as settled until that metric has been shown to credit a positive control.
- constraint Genome-wide screens, where the average perturbation moves only a few genes, are where error metrics are least able to register a model's skill, so leaderboards built on them need recalibration before they rank models.
- capability Benchmark builders now have a per-dataset check on any metric: one that cannot rank the interpolated duplicate above the mean baseline is unfit to score models on that dataset.
In experimental biology, an assay is trusted only when a positive control run through it comes out positive [8]. Perturbation benchmarks had a negative control, the mean baseline, and typically no positive one [8]. A poor score could then mean a failed model or a metric unable to see success, the authors argue [8]. Their first candidate control predicted half of each perturbation's cells from the other half [9]. Both halves carry the same perturbation, and in Norman19 the split beat the mean as expected [10].
It broke on the genome-wide Replogle22 K562 screen, and the authors traced the failure to signal dilution [11][12]. Norman19 moves about 31 times as many genes per perturbation [1]. When a perturbation changes only a handful of genes, mean squared error mostly tests how well a prediction estimates the thousands of genes that did not change [6]. The authors say the mean baseline wins that test largely because it has more statistical power [13]. The few real changes barely register in the score [6].
Pearson delta fails in a different way. Correlations referenced to control cells are inflated when perturbed cells differ from controls in systematic ways [5]. The mean baseline is an average of perturbed cells, so it inherits whatever shift they share, whichever gene was hit [2][5]. Both artifacts first appeared in this group's own earlier work [7]. Control bias has independent support: Viñas Torné et al. described it as well [5].
With calibrated metrics, the paper reports, deep-learning perturbation models "can outperform" uninformative baselines [4]. The field wants these models for in-silico screens that find therapeutic targets [15]. The thing this doesn't tell you is whether any model is accurate enough for that job. Beating an uninformative baseline is the minimum a screening model has to do [15].
I think the earlier negative verdicts, including that of Ahlmann-Eltze et al. [1], should be read narrowly. Some rest on MSE or Pearson delta in datasets whose perturbations move few genes. In those cases they used metrics that ranked half of each perturbation's own cells no better than a training-set average [11][3]. That makes them weak evidence against the models. The case for the models depends on margins, and the excerpt does not report them or name the models that clear the bar [4].
What to watch
- The paper's model-level results: which deep-learning architectures beat uninformative baselines on the calibrated metrics, and by what margin.
- Whether Ahlmann-Eltze et al. or other benchmark authors re-score their comparisons with calibrated metrics and reach the same reversal.
- Whether predictions that clear calibrated metrics hold up when tested prospectively against new wet-lab perturbation screens.