Science1 distinct publisher2 min readPublished
GOLLuM fits a language model and a Gaussian process under a single probabilistic objective, so the model learns from being confidently wrong. Its authors report matching classical Bayesian optimisation on over 40 per cent fewer experiments.
The Scientist · Science desk

Compiled by The ScientistSomething wrong?How this is made
The key design choice is where the uncertainty lives. Classical Bayesian optimisation puts a Gaussian process on top of descriptors a chemist built by hand, and the process returns a prediction together with a variance, which is what lets a search trade exploration against exploitation [7]. GOLLuM keeps the Gaussian process and feeds it a language model's embeddings instead, then trains the language model through that same probabilistic objective [2]. Under such an objective, confident wrongness is what the loss punishes hardest, and the representation has to move to reduce it: the authors report embeddings rearranging so that experiments with similar outcomes end up near each other, which is how structure in the design space becomes visible to the optimiser [3]. Earlier work pairing language models with Bayesian optimisation kept the two as separate components, and the authors present theirs as the first to fit both under one objective [11].
The doubling claim holds up when you run the numbers. Forty-three per cent against 24 to 25 per cent [6] is a gap of 18 to 19 percentage points, or a relative increase of 72 to 79 per cent [14]. The sample-efficiency figure travels better: matching classical Bayesian optimisation on over 40 per cent fewer experiments [5] turns a 100-run campaign into fewer than 60 runs [15], and the run is the unit a lab budget is denominated in. The abstract leaves unspecified the threshold that makes a reaction high-performing, which base model was wrapped, and what the fitting costs [17].
It is unclear whether those 23 tasks were live bench work or selections replayed against outcomes already measured [16]. That distinction carries more weight than usual, because the asset GOLLuM borrows is precisely the pretrained knowledge that lets language models cross domain boundaries [13], and Buchwald-Hartwig coupling is documented well enough that a model may have read the answers before the benchmark started. A clean test needs a design space whose results were published after the base model stopped training. The full paper may contain one; the abstract does not say.
My read, with that condition attached, is that the objective is the interesting part rather than the model. Wrapping a language model in a Gaussian process leaves its suggestions as fallible as they were, while making that fallibility measurable before anyone weighs out a reagent, which is the property that hallucinated suggestions lack [10]. The wider claim the authors attach, that foundation models can be specialised through richer uncertainty-guided information rather than more data [12], is the one I would want tested somewhere nobody has benchmarked yet.
Ranked by verification strength, evidence, and original report placement.
A paper titled 'Large language models as uncertainty-calibrated optimizers for experimental discovery', published on nature.com, introduces GOLLuM (Gaussian Process Optimized LLMs), a framework for using language models as reliable optimizers guided by natural language.
GOLLuM trains language models through Bayesian objectives, teaching them from experimental outcomes under uncertainty and transforming LLM overconfidence from a flaw into a precise learning signal.
The learning signal reshapes the LLM embeddings so that experiments with similar outcomes cluster together, revealing structure in the design space.
Starting from only ten low-performing experiments, GOLLuM generalizes across 23 tasks in organic synthesis, materials science, process chemistry and molecular design, ranking first on average among all competing methods.
GOLLuM matches traditional Bayesian optimization with over 40% fewer experiments.
GOLLuM nearly doubles the discovery of high-performing Buchwald-Hartwig reactions over expert quantum-chemical descriptors and state-of-the-art LLMs, at 43% versus 24-25%.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 27, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
science
Nature Perspective: autonomous agents need internal states they must keep in range1 distinct publisher
science
MAP swaps drug IDs for a mechanism graph and claims zero-shot single-cell predictions1 distinct publisher
science
A jazz model names the player 94% of the time, and says which part of the playing gave them away3 distinct publishers
science
The archive as a neural network: WIEN-INR bets small models can hold full-spectrum instrument data1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Peer-reviewed, single-voiced
The numbers come from a refereed Nature paper and are unusually specific — 23 tasks, ten seed experiments, 43% against 24–25% — which is better provenance than most claims of this shape get. But it is one document, written by the method's authors, and the excerpt available to us stops before methods: the evaluation protocol, the 'high-performing' cutoff and the base model all sit outside what a reader can check.
Nothing past publication day
A paper appearing is not a method being used. We have no lab running GOLLuM, no code or model release, no vendor picking it up, and not even a statement of which base language model someone else would have to fine-tune to try it. Until one of those surfaces, any adoption figure would be invention.
Paradigm talk over benchmark proof
The measurements are modest in their language; the framing around them is not. 'First framework', 'a different paradigm for specializing foundation models', overconfidence 'transformed' from flaw into signal — that is a claim about the future of foundation-model training resting on a benchmark sweep whose provenance we cannot see. Discount the frame, keep the percentages, and the gap is small rather than damning.
The people who named it are the only witnesses
The authors chose the 23 tasks, chose the baselines, coined the acronym and declared the field's first. None of that makes the result wrong; all of it means the framing and the evidence share a single author. What our coverage does not contain — funding, commercial ties, competing groups' responses — would move this number in either direction.
Good provenance, unopened method
Confidence sits mid-range for a specific reason: peer review at Nature is real signal, and the claims are quantitative rather than vibes, but everything traces to one document whose methods section we do not have. A single independent re-run of the Buchwald–Hartwig comparison, or a clear statement that the 23 tasks were physical rather than replayed, would move this substantially.