Skip to content

ScienceNot yet confirmed elsewhere1 publisher2 min readPublished

Active learning chose the lab experiments that taught a model to predict C-H borylation sites

A research team used an active-learning loop to choose its C-H borylation experiments, then pinpointed the correct site in every test on unseen substrates. The work treats a shortage of experimental data as the real limit on predicting reactions.

The Scientist · Science desk

How we use AISend a correction

Photograph accompanying Active learning chose the lab experiments that taught a model to predict C-H borylation sites
Photo: nature.com

What happened

  • The lab campaign produced a diverse benchmark dataset for the C-H borylation reaction.
  • Adding self-supervised auxiliary tasks improved the models' predictions of both whether a reaction works and where on the molecule it happens.
  • The substrates used in the prospective test carried challenging N-heteroaryl motifs.

Why it matters

  • constraint If the binding limit is a shortage of data, a bigger network will not move the result. The leverage is in choosing which experiment to run next.
  • decision Scoring the surrogate against millions of candidate molecules every round made deep neural ensembles too costly, so the loop's selector is a cheaper tree-based model.
  • capability Because the models call the reaction site on substrates they never saw, a chemist can ask where a new drug-like molecule will borylate before committing lab time.

The workflow splits the work between two models. A tree-based ensemble picks which experiments to run [1]. The predictions come from geometric graph neural networks, which read a molecule as a 3D object with rotational, translational and permutational symmetry built into the network [11]. Between the two, the priority changes: collecting data rewards fast iteration, while a finished dataset rewards raw predictive accuracy [10].

On the central result, the abstract is thin. It does not say how many substrates the prospective test covered, or give a numeric baseline to beat [5]. The self-supervised gain is described only as "consistently" better [4].

Underneath the engineering is a claim about where the bottleneck sits. The authors write that identifying and optimizing drug molecules is "increasingly limited by the cost and speed of experimental data generation rather than by molecular design algorithms" [12]. So the effort goes into choosing the next experiment well. It matters upstream in the design software too: a generated molecule is only worth proposing if it can be made. Today's reaction models are held back by scarce data and by how hard regioselectivity is to predict [14].

Two limits sit on the result. It is demonstrated on a single reaction class, C-H borylation [2], so it does not show how the loop behaves where selectivity follows other rules. And the stated goal is to map the whole reaction landscape, negative results included, so the models generalize to substrates no one has run yet [13].

What to watch

  • Whether the loop and its models hold up on reaction classes beyond C-H borylation, where selectivity follows other rules.
  • A published dataset size and a numeric baseline, so the perfect prospective score can be sized against alternatives.
  • Whether chemists wire the regioselectivity predictor into generative design pipelines to screen proposed molecules for synthesizability.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence52
Adoption
Insufficient
Hype gap+12
Incentives55
Confidence48
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    An active learning strategy, driven by a tree-based ensemble, guides the acquisition of the most informative laboratory experiments.

    ReportedSupportedView cited source
  2. [2]

    The experimental campaign yielded a diverse benchmark dataset for C-H borylation.

    ReportedSupportedView cited source
  3. [3]

    The bespoke dataset was used to train and evaluate geometric graph neural networks.

    ReportedSupportedView cited source

Sources

1 independent publisher whose own reporting we read for this story.

  1. nature.com

    1 article · October 8, 2026

    Advancing chemical reaction prediction in data-scarce drug discovery with active and geometric deep learning

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Topics

Entities

Loading related stories