Science1 distinct publisher2 min readUpdated
A Monell-co-authored DREAM challenge turned odor-mixture similarity into a public machine learning task. The winning ensemble's error and its correlation tell different stories about how close this is.
The Scientist · Science desk
Compiled by The ScientistSomething wrong?How this is made
A median RMSE of 0.08 reads like a solved problem. On a scale running from 0 for indistinguishable to 1 for most distinct, the median miss is eight percent of the entire perceptual range [10][6][1]. The Pearson correlation on the same test set is 0.57, which squares to about 0.32, so under a third of the variation in what people actually judged is accounted for [11][2]. Both numbers can hold at once. Small absolute errors alongside mediocre rank agreement are what you get when most pairs sit in a narrow band and a model predicting near the middle is rarely badly wrong. Which number matters depends on the job. A tool meant to flag a copied detergent scent needs the ordering. A substitution model inside a formulation pipeline may be satisfied with the absolute error.
The more useful finding is what the winning models fed on. According to Mainland, they reached for semantic descriptors such as "fruity and sweet" rather than chemical features like ester class or molecular weight [14]. When a journal reviewer made the group strip the semantic features out, prediction became much harder [15]. The map, then, is not built from chemistry. It sits on a layer of human verbal labels attached to the single components, with the mixture arithmetic on top. Mainland's own summary is that if you can predict what one component smells like, mixtures follow reasonably well [13]. That is a genuine result, and it also relocates the hard part back to single-molecule prediction, which is where IBM's 2015 DREAM challenge left it [3].
The benchmark's weight is worth sizing before anyone leans on it. There are 507 pair measurements across 731 unique mixtures, which is 0.69 measurements per mixture [5][3], and the set is a harmonisation of six earlier datasets drawn from three studies, so its measurement conventions are inherited rather than fresh [5]. Scoring rests on 96 pairs in total: 46 in the hidden test set and 50 in the independent validation [7][9][4]. Set that against everyday odors made of dozens or hundreds of molecules [4] and the sampled slice of the combinatorial space is very small.
None of which undercuts the thing that actually changed. Before this, a claim about digitising scent could be evaluated only by reading the description of it. Now there is a validated metric, a public target and a named set of pairs [2], and the score to beat is 0.08 and 0.57 [10][11]. The field expected mixtures to be much harder than single molecules and, per Mainland, they were not [12]. That expectation was doing quite a lot of work as an excuse.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Research co-authored by scientists at the Monell Chemical Senses Center created a means of using machine learning to distinguish among scents and how they relate to one another; it was published in the Proceedings of the National Academy of Sciences.
In 2015, IBM ran an open investigator DREAM challenge to take a single odor molecule and predict what it smells like based on its chemical structure.
Most odors encountered in day-to-day life are complex mixtures of dozens or hundreds of molecules, and Mainland says mixtures must be understood in order to digitize anything.
The organizers standardized and compiled six data sets of odor-similarity measurements from three different studies into a new data set comprising 168 unique single molecules, 731 unique mixtures and 507 mixture-pair measurements.
Mixture-pair distances were mapped onto a continuous perceptual scale from 0 (indistinguishable) to 1 (most distinct).
Over a three-month period, 26 teams competed to predict how similar paired scents would be on a hidden test set of 46 mixture pairs, and the competition ended in a four-way tie.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Peer-reviewed and numerically specific, but single-source and small-scale
The underlying work is published in PNAS with a DOI, and the article reports unusually concrete specifics: data set counts, the bounded target scale, challenge structure, ensemble construction, and two headline metrics. Against that, the entire cluster rests on one publisher relaying study-team quotes, the evaluation covers only 96 mixture pairs, and the correlation figure implies most rating variance is unexplained. The forward-looking claims (digital olfaction foundation, diagnostic and industrial applications, third-challenge findings) carry no independent corroboration.
Research-community uptake only; no deployment
Observable uptake is confined to the scientific community: 26 teams entered the mixture challenge, the benchmark and ensemble were published, and a third challenge has been run with results pending. Nothing in the supplied material shows a product, company, sensor platform, or operational user adopting the metric, and the article does not state whether the data set or models are publicly available for reuse.
Accuracy framing outruns the reported correlation and sample size
The article calls the model 'quite accurate' and frames the result as laying a foundation for digital olfaction, while the same text reports a test-set Pearson of 0.57 — roughly 32 percent of variance explained — on a total of 96 mixture pairs. The low median RMSE is partly an artifact of a bounded 0-1 scale where many pairs cluster, so it flatters the model relative to the correlation. Diagnostic, flavor and trademarking applications are presented as near-at-hand without any demonstration. The overstatement is moderate rather than severe: the numbers themselves are disclosed rather than hidden.
Study team is the sole narrator; outlet reprints institutional framing
Every interpretive statement — that the metric is validated, that mixtures turned out easy, that industry could work more efficiently, that a third challenge confirms the averaging result — comes from a co-author of the paper whose institution benefits from visibility for its digital-olfaction program. The publisher's format is science-release aggregation, so no adversarial questioning is applied. There is no evidence of undisclosed commercial funding or vendor sponsorship in the supplied material, which keeps this short of the top of the scale.
Core facts firm, interpretation weakly grounded
Confidence in the mechanical facts — publication, data set composition, challenge format, the two metrics — is high because they are stated precisely and are internally consistent. Confidence in the significance claims is low: one publisher, one interested narrator, no independent replication, no adoption beyond the research community, and a headline metric pair that points in two directions. That mix supports a middling overall confidence.
product
Geologic hydrogen's firmest numbers come from mine vents, not reservoirs1 distinct publisher
science
IBM's bottleneck was cold volume, and its answer is an 8-foot box you can bolt to another one1 distinct publisher
science
OX Security says MCP command execution is a design choice, so server owners own the risk1 distinct publisher
product
Inference cost is now a storage and power problem, and it prices differently per workload1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
phys.org
1 article · August 21, 2026