Science1 publisher3 min readPublished
Two-step model predicts chromatography retention times for systems it was never trained on
Researchers' 2-step method predicts small-molecule liquid chromatography retention times on a new system using no training data from that system. The best current predictors must first be fine-tuned on hundreds of standards per instrument, and a model that transfers would let labs skip that work.
The Scientist · Science desk

What happened
- The first step is a machine-learning model that predicts each compound's retention order index under the given conditions, and a second step converts those indices to absolute times.
- The design relies on retention order being far more conserved than retention time, which can change massively even under nominally identical conditions.
- The authors also ran a systematic study of which chromatographic conditions cause notable changes in retention order.
Compiled by The ScientistSomething wrong?How this is made
Why it matters
- cost If the transfer claim holds, a lab would no longer have to buy and run hundreds of authentic standards on each instrument before it can get a retention time prediction.
- capability A lab could compare measured retention times against predictions for candidate structures on its own column, using retention as identity evidence, without first building a training set there.
- constraint The method is built for reversed-phase separations, so labs that depend on other separation modes cannot use it as described.
Retention time depends on the whole setup: the column, the eluents and their composition, the compound and the rest of the LC equipment [8]. Four decades of work and many hundreds of papers have still not put small-molecule prediction into everyday use [8]. A typical study measures at most a few hundred compounds on one fixed platform. The model it trains then fails on new conditions, or on molecules structurally unlike its training set [9]. The authors' bet is that order carries across systems, and that the system-specific part can be handled in a separate conversion to minutes [4].
Gas chromatography already normalises across conditions with retention indices built from a set of standards, and the idea has recently been proposed for LC [12]. The authors point out its weak spot. Elution order can change when conditions change [12]. Their order index is predicted with the chromatographic conditions as an input [3], and the reordering study shows which conditions are most likely to reshuffle it [6].
I would read the comparison most closely. According to the authors, 2-step beats existing methods that were trained on the target dataset [5]. That is a demanding control, because the baselines saw data from the system under test and 2-step did not [4]. The current leaders use transfer learning. They pretrain on a large dataset with arbitrary conditions, then fine-tune on hundreds of authentic standards measured on the target system [11]. Matching them without that step would be a real result. The abstract and introduction do not include error figures, the number of systems tested or the size of the margin. They also do not describe what the second step uses to tie an index to minutes on a new system. That detail decides how much work a lab still does before the model is usable.
For a working lab, the problem is variation between instruments. A fine-tuned model cannot be reused on another system, even a nominally identical one [11]. The authors themselves write that retention times can shift massively between such systems [2]. A good score on pooled public datasets shows a method ranks well. The figure a lab needs is its error on a column, in another lab, that contributed nothing to training. I think the decomposition is sound, provided the order model holds up on column chemistries it has not seen.
The scope is narrow on purpose. The method covers reversed-phase chromatography [3] of molecules below 1,500 Da. For these, prediction has proved far harder than for peptides or oligonucleotides [7]. The closest earlier multi-condition method, MultiConditionRT, accounted for eluent composition but did not handle different columns beyond separation mode [10].
What to watch
- Error figures in the full paper: the 2-step prediction error on held-out systems against fine-tuned baselines, and how many columns and systems that comparison spans.
- What the retention-time mapping step needs from a new system, such as the gradient program, dead time or a few reference compounds.
- Independent groups testing 2-step on their own reversed-phase setups, or any extension to other separation modes.