Skip to content

Science1 publisher2 min readPublished

resolVI corrects misassigned RNA molecules downstream of any segmentation tool

A paper published by Nature treats stray RNA molecules as a characterizable artifact and corrects them statistically after cells have been drawn. The abstract reports better cell-state separation.

The Scientist · Science desk

Illustration accompanying resolVI corrects misassigned RNA molecules downstream of any segmentation tool

What happened

  • resolVI produces a probabilistic representation of spatial transcriptomics data that corrects for molecules assigned to the wrong cell, for batch effects and for other nuisance factors.
  • The model uses autoencoding variational Bayes and returns artifact-corrected probabilistic estimates of both a low-dimensional representation of each cell and its gene expression profile.
  • The authors trace misassignment to RNA leaking during tissue handling and to cells overlapping along the axis perpendicular to the imaged plane, which is profiled across only tens of microns.
  • They report that the correction improves separation of cell states, detection of subtle expression changes across space, and integrated analysis of several datasets together.
  • The code is released as open source software within the scvi-tools package.

Compiled by The ScientistSomething wrong?How this is made

Why it matters

  • decision A group unhappy with its cell-state calls can weigh adding a correction step against buying or building a new segmentation pipeline.
  • capability Because the model takes an existing per-cell expression table as input, archived datasets can be reanalysed without re-running segmentation or collecting new tissue.
  • constraint With no effect sizes in the abstract, a team cannot yet judge how much of a newly separated cell state is recovered signal and how much is the model's assumption about diffusion.

High-resolution spatial transcriptomics comes in two families. One sequences RNA captured on submicrometer barcoded spots; the other images individually tagged molecules. Both need the tissue plane divided into cell-sized regions before there are cells to analyse [14].

Most of the effort has gone into drawing those regions better. The early algorithms traced nuclei and membranes in images of the tissue [7]. They run into trouble where cells are densely packed, where the staining signal is weak and where shapes are irregular [8]. Newer methods added the observed RNA molecules as a second source of information to guide the boundary. The authors report that this coupling largely improved accuracy while the same problems persist, and they demonstrate it in the paper [10].

resolVI leaves the boundaries alone. Its input is the initial per-cell estimate of gene expression that an upstream pipeline already produced [1]. The authors write that the basis of the model is the realization that diffusion of signal is an artifact that should be characterized and then controlled for [13].

Misassignment shows up in the data as biological nonsense, a cell that appears to express marker genes from two different lineages [9]. In the abstract the authors wrote of a persistent problem: "the incorrect assignment of molecules to cells, which limits many current applications to the level of a priori-defined cell subsets and complicates the discovery of novel cell states" [6].

The abstract gives no effect sizes, tissues or baseline pipelines for its improvement claims [15]. Two measurements matter for a lab deciding whether to add the step. One is how often a corrected profile flips a cell-type call. The other is how often the correction erases a difference between neighbouring cells that was real.

One class of error sits outside what the correction can reach. The authors note that inferred segments may get the number or the shapes of the cells wrong [12]. A region that merged two cells arrives in the model as one cell, because what the model is handed is the per-cell table [1].

What to watch

  • The full paper's benchmark tables: which tissues, which segmentation baselines, and how large the reported gains are.
  • Whether independent groups reproduce the cell-state gains on their own data using resolVI as shipped in scvi-tools.
  • Whether any evaluation reports the false-positive side: cell states that appear only after correction.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories