Skip to content

Science1 publisher3 min readPublished

Böcker's Jena group predicts chromatography retention times by ranking molecules first

Sebastian Böcker's Jena group reports in Nature Methods a tool that predicts chromatography retention times on new lab systems without retraining. Whether that makes retention time a dependable filter for identifying unknown molecules depends on how large its prediction errors are.

The Scientist · Science desk

Photograph accompanying Böcker's Jena group predicts chromatography retention times by ranking molecules first
Photo: nature.com

What happened

  • Earlier prediction models usually had to be trained or fine-tuned on the same instrument setup, after the lab had measured many standard substances.
  • The method's first step is a machine-learning model that gives each molecule a retention order index, its expected place in the elution sequence.
  • A second step turns that index into actual retention times using a few known reference compounds.
  • The tool, called 2-step, is available as a software package and as a web application.

Compiled by The ScientistSomething wrong?How this is made

Why it matters

  • cost Recalibrating after a hardware change would cost a lab only a few reference runs, small enough to repeat after every column or tubing swap.
  • decision Labs considering 2-step as an identification filter will need to check its error on their own columns before letting it discard candidate structures.
  • constraint Groups that separate molecules with modes other than reversed-phase, the mode the method is built for, get no help from it as published.
  • precedent If analysis software adopts it, as the researchers propose, retention checks would run automatically on LC-MS identifications without a chemist setting them up.

Retention time is the moment a molecule leaves a liquid-chromatography column, and it gives a clue to which substance is present [16]. It has been hard to predict because it depends on the instrument as much as on the molecule. "The difficulty lies in the fact that retention times depend heavily on the experimental conditions: on the column used, the solvent, the gradient, the pH, the temperature and even on seemingly minor technical changes," Böcker said [3]. He gave an example: "If, for example, a tube in the apparatus is replaced, or if a new tube is slightly longer than the old one, the measured times can shift significantly." [4]

The new method separates those two dependencies. Its learned step predicts only where a molecule falls in the elution order [7]. A replaced tube or a new column then has to be absorbed by the second step, a conversion to minutes fitted on a few reference compounds [8]. The paper's title states the premise: "Times are changing but order matters" [15]. I think this is a well-chosen split. Order is the quantity a model trained on chemical structures can plausibly carry from one lab to another. A few reference runs are what a lab can afford to measure on its own instrument.

The work comes from Böcker's group at Friedrich Schiller University Jena, with partners at Helmholtz Zentrum München and the Technical University of Munich [1]. Fleming Kretschmer, a co-first author who worked on the method as part of his doctorate [11], said: "A key finding of our work is that our method outperforms other approaches that, unlike ours, first have to be trained extensively on the target system." [9] He added: "Our approach therefore enables precise predictions out of the box, even for new systems." [10] If it holds up, that is a notable result, because the competing models had trained on data from the very system being tested [9]. The press account of the study does not report error sizes, the margin over those models, or how many reference compounds the conversion step needs.

Those figures decide whether retention time becomes a working identification filter. Labs in environmental analysis, food chemistry and pharmaceutical research keep asking whether a measured retention time matches the presumed chemical structure [14]. A prediction answers that only when its error is smaller than the gap between the candidate structures a lab is choosing among. Once the error exceeds that gap, the check passes the wrong candidate as readily as the right one. The thing this doesn't tell you is how many wrong candidates 2-step would strike from a real shortlist, on a column that has just had a tube replaced.

What to watch

  • The paper's reported prediction errors in minutes for each test system, and how many reference compounds the calibration step needs.
  • Independent labs running 2-step on their own reversed-phase setups, especially right after a column or tubing change.
  • Tests that count how often a 2-step retention check removes wrong candidate structures from real identification shortlists.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories