Science1 distinct publisher3 min readPublished
Rather than wait for experimental libraries to fill, the Cornell and BTI tool indexes a chemical space roughly a thousand times wider. What a query returns is a candidate structure rather than a confirmed one.
The Scientist · Science desk

Compiled by The ScientistSomething wrong?How this is made
Divide the index by its contents and you get about eight predicted spectra for every molecule: more than 800 million spectra covering more than 100 million compounds [14]. The write-up does not say what the other seven capture, and it does not report how often a retrieved match is the right one [19]. What it does establish is that the spectra were generated before anyone searched them and then filed into a browsable resource [3], so the expensive model inference was paid once, by the authors, instead of per query by every lab.
Run the thousandfold claim backwards and the incumbent looks thin. A thousandfold expansion of a space containing more than 100 million molecules [2][5] puts existing experimental coverage on the order of 100,000 compounds, which is about 0.1% of that space [15]. That is consistent with the paper's own accounting, which puts collective library coverage at under 1% of known compounds [4]. Spectra without a close library relative usually just stay unannotated [18], which is the mechanical reason the 80% figure is as large as it is [1].
The design choice that matters is what comes out alongside the predicted spectrum. DeepMS2Reasoner enumerates physically plausible fragmentation steps using symbolic chemical rules, uses a neural network to assign each step a likelihood, and returns an annotated account of how the molecule came apart [7]. A library hit gives a chemist a similarity number to trust or distrust. A fragmentation route gives them something to disagree with, which is a different and more useful kind of output.
The demonstration's denominator deserves attention. Several thousand chemical features separated germ-free mice from mice with normal gut microbiota, and most of the abundant ones defeated standard identification [10]. The team queried the index with 111 of them, selected as the most abundant unidentified microbiota-dependent compounds [11]. Roughly a third returned close predicted matches with related structural candidates [11], which is about 37 compounds; the other 74 or so came back only as molecular neighborhoods of structurally related compounds [12][16]. Two compounds went further, with spectra pointing to polyamine derivatives, according to Frank Schroeder of the Boyce Thompson Institute and Cornell [13][8].
So the tested cases were a selection of the loudest signals in the data set, the most abundant unidentified compounds rather than a random draw from the unannotated pile, and the account does not say how retrieval behaves on fainter features [11][19]. Nor does the ranking change the cost of being sure: by the authors' own description, pinning down a single unknown structure runs days to months of iterative analysis and experimental validation [6]. A neighborhood of candidates shortens the reasoning, not the bench work.
Sizing how much of that 80% AIMe actually recovers requires retrieval precision measured against spectra whose structures are already known. The announcement reports coverage [19].
Ranked by verification strength, evidence, and original report placement.
Schroeder says two compounds became a case study, both producing spectra that suggested they were polyamine derivatives.
More than 80% of compounds detected in a typical biological sample cannot be matched to any known structure using current methods.
AIMe (AI Molecule Explorer), from the Boyce Thompson Institute and Cornell University, uses neuro-symbolic AI to predict, organize and search the mass spectra of more than 100 million known small organic molecules.
AIMe generated more than 800 million predicted spectra covering essentially all known small organic molecules in PubChem, organized into a searchable resource called MS2KOSMOS.
Available experimental reference libraries collectively cover fewer than 1% of known compounds.
MS2KOSMOS represents roughly a thousandfold expansion of searchable chemical space relative to existing experimental libraries.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · September 4, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
science
A nematode's pheromone primes plant defense genes by loosening their chromatin1 distinct publisher
science
No country is on track to cut food system emissions by 2030, first goal-referenced audit finds1 distinct publisher
science
Sucralose and stevia reshaped mouse gut chemistry into generations that never drank them1 distinct publisher
science
Cornell's Manhattan CO2 twin suggests the hard part is plumbing, not sensing1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Single account, pre-peer-review
Every number traces to one telling by the people who built the tool, published before review. The chemistry is specific enough to be checked later: a named model, a named resource, a linear putrescine derivative and a macrocyclic polyamine each confirmed by synthesising an authentic standard, and that macrocycle then found in 57 of 99 human fecal samples in a public repository. What the account never supplies is the figure that would let a reader judge the resource rather than the case study, namely how often retrieval from 800 million predicted spectra returns the right structure when the answer is already known.
Free to use, used only in-house
The tool is live and the preprint is posted, which is more than many announcements at this stage offer. The recorded use, though, is all the group's own: the germ-free mouse comparison, and a pass over more than 7 million GNPS spectral clusters whose outcome the account declines to state. No external lab, collaborator or reviewer is quoted using MS2KOSMOS, and since the code arrives with publication there is no reproduction to point to.
A wider index described as identification
Mapping the hidden universe of small molecules and a thousandfold expansion of chemical space are large phrases for what the demonstration delivers: candidate structures for about 37 of 111 abundant unknowns, neighborhoods for the remaining seventy-odd, and synthesis required to settle each of the two compounds actually pursued. The predicted library is genuinely large and the interpretable fragmentation map is a real advantage over a bare similarity score. The overreach is in letting a searchable index of predictions stand in for annotation.
Institutional release, lightly syndicated
This is the Boyce Thompson and Cornell communications account of Boyce Thompson and Cornell work, timed to a preprint and to a tool the group wants used: institutional voices only, a link to the running service, code held until the paper lands. phys.org carries such releases with little editing, so which numbers lead, 800 million and thousandfold rather than 37 of 111, was decided by the authors. Nothing here suggests a commercial interest; the pull is toward uptake and citation.
Build verified, performance unmeasured
It's clear what was built and what it covers, and reaching it is no mystery either, and one structurally unusual macrocyclic polyamine has synthesis behind it. How MS2KOSMOS behaves at scale is not assessable from what is here: no accuracy figure against known spectra, and no user outside the group. Peer review will either supply that or it will not.