Skip to content

Science1 publisher2 min readPublished

Molecular Community Network links nearly 95% of mass-spec molecules to at least one neighbor

Vladimir Boginski's Molecular Community Network connects nearly 95% of mass-spec molecules, so unknowns can take candidate identities from neighbors. Its use for biomarkers depends on how often those predictions survive lab checks, as one class of gut-microbe bile acids has.

The Scientist · Science desk

Photograph accompanying Molecular Community Network links nearly 95% of mass-spec molecules to at least one neighbor
Photo: ucf.edu

What happened

  • Vladimir Boginski and an interdisciplinary team published the Molecular Community Network, a method for organizing untargeted metabolomics data, in Cell Reports Methods.
  • Boginski estimates that up to 90% of observed molecular space is dark matter, detectable by mass spectrometry but matched to no known structure.
  • The method groups the whole network into communities before keeping its strongest links, leaving nearly 95% of molecules connected and assigned to a community.
  • It predicted a new class of gut-microbe bile acids, one of which appears only in young infants, and the team then confirmed the class in the lab.

Compiled by The ScientistSomething wrong?How this is made

Why it matters

  • cost The first pass at naming an unknown now needs computing time on spectra already deposited in public repositories, and no new samples.
  • constraint Because every propagated identity is a prediction, the rate of confirmed discoveries is capped by how many candidates labs can check at the bench.
  • contradiction Boginski uses 8.4 million as a count of spectra in one quote and of molecules in another, so the size of the mapped molecular space is less certain than either figure suggests.

Classical molecular networking draws an edge between two molecules only when the similarity score of their mass spectra clears a preset threshold [4]. Relatives that score just under the cutoff land in separate pieces, so one chemical family can be split across the map [4]. That matters for unknowns, because identification by network works through known neighbors [7].

The Molecular Community Network starts from the whole network. Its algorithm partitions it into communities, groups with many strong links inside and fewer between them, and then keeps the strongest connections so the communities stay joined [5]. "This approach allows us to have almost every molecule in the network linked to at least one neighbor," said Vladimir Boginski, who led the work with first author Elizabeth A. Coler and an interdisciplinary team [1][2]. "Moreover, these links are typically between molecules from similar molecular families." [6] An unknown that shares a community with identified molecules can then be given a predicted identity from them. The field calls this step annotation propagation [7].

If Boginski's upper estimate of 90% dark matter held across the roughly 8.4 million spectra in public repositories [3], up to about 7.6 million of them would be detectable but unmatched to any known structure [12][1]. With nearly 95% of molecules now in communities [8], about 420,000 or slightly more would still sit outside one [2].

The thing this doesn't tell you is how often a propagated identity is right. Connection gives an unknown what Boginski called "a much wider and richer search space for molecular discovery" [8]. The public account of the paper does not include an accuracy rate for predicted identities, or a connectivity figure for threshold networking on the same data.

The bile-acid class is the one direct test in that account. The network made the prediction first, and lab work confirmed it afterward [9]. That order is the right design, because the prediction could have failed at the bench.

Boginski puts the payoff in biomarker work. "Biomarker discovery often stalls at the point where a molecule that differentiates sick from healthy is detected by its mass spectrum, but it may not be known what this molecule is," he said [10]. "MCN can potentially help convert some of those 'dead ends' into identifiable molecules," he said [11]. I think "some" is the accurate word, since the published account describes one confirmed class so far [9].

What to watch

  • An accuracy rate for annotation propagation, measured by hiding the identities of known molecules and checking what the network predicts for them.
  • Connectivity figures for threshold-based molecular networking on the same repository data, which would size the gain.
  • A disease biomarker named through the network and confirmed by a lab outside Boginski's team.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories