Science1 distinct publisher2 min readPublished
Queries that once meant downloading petabytes of raw files now run in seconds against a reference-free index, which also makes isoform-level and off-reference variant questions askable rather than merely faster.
The Scientist · Science desk

Compiled by The ScientistSomething wrong?How this is made
Reference-free is the load-bearing phrase in the paper's title [13]. Existing platforms map their data onto a reference genome, which is a composite of a few individual human DNA samples, so RNA produced from a rare or unique gene variant may leave little trace in the index [8]. They also register isoforms as one entity per gene rather than indexing each individually, which is why you cannot currently ask whether one isoform is expressed in one cell type [7]. The Berlin group reprocessed public repository data and keyed it on nucleotide sequence instead [14], turning both dead ends into ordinary lookups.
The speed figure deserves its own units. The old route is described as taking at least several days across data from thousands of experiments [3]; read "several" as three and that is 259,200 seconds [16], against a stated query time of seconds [4]. I would not turn that into a ratio and put it on a slide. The source gives no distribution of query times and no workload definition for the multi-day figure, and the two numbers measure different objects: one is wall-clock for a human-supervised pipeline, the other a hosted lookup.
The thing this doesn't tell you is how often a sequence query returns what an aligned pipeline would have returned. The announcement offers "millions of cells" as scale [4] and reports no count of indexed datasets and no precision or recall against alignment-based results [17]. It also does not describe how cell-type annotations from different sources were harmonised [18], and cell type is the axis most of these queries resolve against.
Two features widen the question space more than the latency does. The index covers spatial data, so a hit can be placed within a tissue section [9], and it carries sequence from bacteria, viruses and fungi [10]. That combination is what makes the announcement's framing scenario, doctors asking which cells drove a tumour and whether pathogens contributed [15], a thing you can type into a box. Typing it is not attribution. A microbial sequence sitting among tumour cells is co-occurrence, and nothing in a retrieval index establishes which came first.
My read, conditional on the Nature paper reporting sensible retrieval accuracy: the choice of key matters more than the seconds, because a class of RNA questions becomes cheap enough to ask on a hunch, and that is usually how a method changes practice rather than benchmarks.
Ranked by verification strength, evidence, and original report placement.
Researchers at the Berlin Institute of Medical Systems Biology of the Max Delbruck Center (MDC-BIMSB) present Malva, a search engine for single-cell RNA data, in a study in Nature.
Anyone wishing to mine existing single-cell data for a DNA or RNA sequence of interest would need to download and reprocess petabytes of raw files, a task no single lab has the capacity to do, and figure out how to standardize data from different sources.
Because of the way single-cell data is indexed, information about RNA isoforms, meaning multiple RNA variants encoded by the same gene, is extremely limited.
Co-first author Nikos Karaiskos says other platforms cannot answer questions about RNA biology because RNA isoforms are not individually indexed to a reference gene but treated as a single entity, making it impossible to distinguish whether a particular isoform is expressed in a specific cell type.
Malva also indexes spatial data, so researchers can find where in a tissue section a particular RNA is located.
The Malva platform includes sequence information from multiple species, including bacteria, viruses and fungi, so researchers can study how these microorganisms affect human cells and cause disease.
Distinct publishers with included, body-backed reporting in this cluster.
phys.org
1 article · August 28, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
science
A pole-on magnetar hands vacuum birefringence its first astrophysical candidate1 distinct publisher
science
A phage kinase with no target list: EMBL finds one enzyme that breaks several bacterial defences1 distinct publisher
science
HIPAA Covers Less Than You Think, And "Anonymized" Is Not A Legal Shield1 distinct publisher
product
LLNL closes a 20 percent gap in diamond melting, and stakes a fusion gain claim on it1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One release, one DOI, no numbers
A peer-reviewed Nature paper sits under this, which is more than most tool launches can say — but our coverage reaches it only through the Max Delbrück Center's own announcement, relayed by Phys.org. The design facts survive that: sequence-level indexing, spatial and multi-species content, the ingestion loop. The comparative facts do not, because the release carries no index size, no accuracy check against alignment, and no named rival to be better than.
Open to all, used by no one named
Free access plus a live ingestion pipeline is a real starting position, not traction. Phys.org reports no user count, no lab outside the Rajewsky group querying the index, and no downstream result obtained with it; the clinical scenario that opens the piece is explicitly imagined. The commercialisation signal points forward, not to uptake — a startup in formation and a patent still pending.
Google comparison outruns the numbers
Strip the framing and something genuinely new remains: questions about isoforms and off-reference variants become askable, not merely faster, and that distinction is the strongest thing in the reporting. The framing does not stop there. "Like Google did for the internet 30 years ago", "previously impossible questions", seconds against several days, and a much more diverse dataset all arrive without a single measurement, and the tailored cancer treatment is a scenario the reader is asked to imagine in the first sentence. The overreach is in the multipliers and the clinic, not in the premise.
Patent pending, company forming
The people quoted describing Malva as unprecedented are the people patenting it and forming a company to sell it, and the text says so plainly in its penultimate paragraph. The vehicle is an institutional press release, whose job is to make a Nature paper legible and its spin-out fundable; Phys.org adds distribution, not scrutiny. Disclosure is to the release's credit — but it also explains why the superlatives are unmeasured and why no competitor is named.
Confident on design, not on magnitude
Two things are firm: the paper exists, and the index is keyed on sequence rather than reference-aligned genes. Almost everything a reader would act on — how fast, how big, how accurate, how much more diverse, on what terms after commercialisation — comes from one interested party with no figures and no second account to check it against.