ScienceNot yet confirmed elsewhere1 publisher2 min readPublished
Active learning chose the lab experiments that taught a model to predict C-H borylation sites
A research team used an active-learning loop to choose its C-H borylation experiments, then pinpointed the correct site in every test on unseen substrates. The work treats a shortage of experimental data as the real limit on predicting reactions.
The Scientist · Science desk

What happened
- The lab campaign produced a diverse benchmark dataset for the C-H borylation reaction.
- Adding self-supervised auxiliary tasks improved the models' predictions of both whether a reaction works and where on the molecule it happens.
- The substrates used in the prospective test carried challenging N-heteroaryl motifs.
Why it matters
- constraint If the binding limit is a shortage of data, a bigger network will not move the result. The leverage is in choosing which experiment to run next.
- decision Scoring the surrogate against millions of candidate molecules every round made deep neural ensembles too costly, so the loop's selector is a cheaper tree-based model.
- capability Because the models call the reaction site on substrates they never saw, a chemist can ask where a new drug-like molecule will borylate before committing lab time.
The workflow splits the work between two models. A tree-based ensemble picks which experiments to run [1]. The predictions come from geometric graph neural networks, which read a molecule as a 3D object with rotational, translational and permutational symmetry built into the network [11]. Between the two, the priority changes: collecting data rewards fast iteration, while a finished dataset rewards raw predictive accuracy [10].
On the central result, the abstract is thin. It does not say how many substrates the prospective test covered, or give a numeric baseline to beat [5]. The self-supervised gain is described only as "consistently" better [4].
Underneath the engineering is a claim about where the bottleneck sits. The authors write that identifying and optimizing drug molecules is "increasingly limited by the cost and speed of experimental data generation rather than by molecular design algorithms" [12]. So the effort goes into choosing the next experiment well. It matters upstream in the design software too: a generated molecule is only worth proposing if it can be made. Today's reaction models are held back by scarce data and by how hard regioselectivity is to predict [14].
Two limits sit on the result. It is demonstrated on a single reaction class, C-H borylation [2], so it does not show how the loop behaves where selectivity follows other rules. And the stated goal is to map the whole reaction landscape, negative results included, so the models generalize to substrates no one has run yet [13].
What to watch
- Whether the loop and its models hold up on reaction classes beyond C-H borylation, where selectivity follows other rules.
- A published dataset size and a numeric baseline, so the perfect prospective score can be sized against alternatives.
- Whether chemists wire the regioselectivity predictor into generative design pipelines to screen proposed molecules for synthesizability.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence52
- Adoption
- Insufficient
- Hype gap+12
- Incentives55
- Confidence48
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
An active learning strategy, driven by a tree-based ensemble, guides the acquisition of the most informative laboratory experiments.
- [2]
The experimental campaign yielded a diverse benchmark dataset for C-H borylation.
- [3]
The bespoke dataset was used to train and evaluate geometric graph neural networks.
- [4]
Augmenting the symmetry-aware models with self-supervised auxiliary tasks consistently improved performance for both reaction outcome prediction and regioselectivity prediction.
- [5]
Prospective tests on unseen substrates pinpointed the correct borylation positions in all cases.
- [6]
The unseen test substrates featured challenging N-heteroaryl motifs.
- [7]
In this setting the surrogate is retrained after every active learning iteration and evaluated on millions of candidate molecules to select the next points for labeling.
- [8]
Deep neural ensembles estimate epistemic uncertainty by training several independent models and are among the strongest baselines for active learning, but their accuracy carries a steep computational cost.
- [9]
Tree-based methods remain competitive with deep learning at a fraction of the compute budget, making them well suited as an oracle for iterative active learning in data-scarce medicinal chemistry.
- [10]
Given a sufficiently informative dataset, the focus shifts from rapid iteration for data collection to maximizing the predictive performance of production models.
- [11]
Geometric graph neural networks operate directly on 3D molecular graphs, encoding translational, rotational and permutational symmetries through architectural constraints.
- [12]
Rapid identification and optimization of pharmacologically active molecules are increasingly limited by the cost and speed of experimental data generation rather than by molecular design algorithms.
- [13]
In reaction optimization the goal is to map the reaction landscape, including negative outcomes, to build models that generalize to unseen substrates, and it demands atom-level annotations such as regioselectivity that require experimental characterization within a closed loop.
- [14]
The success of generative molecular design depends on the synthetic accessibility of the molecules it creates, and current computational reaction models are limited by a lack of reference data and the difficulty of accurately predicting regioselectivity.
Sources
1 independent publisher whose own reporting we read for this story.
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.
Topics
- AI Drug DiscoveryFollow
- C–H borylationFollow
- Late-stage functionalizationFollow
- Active learningFollow
- Self-supervised learningFollow
- Graph Neural NetworksFollow