Science1 distinct publisher3 min readPublished
Two WashU groups arrive at the same point from opposite ends: the prediction step works, and the corpus of machine-readable synthesis recipes it needs is still being assembled.
The Scientist · Science desk

Compiled by The ScientistSomething wrong?How this is made
Ten viable candidates out of roughly 4,000 curated linker transformations is a hit rate near 0.25%, about one in four hundred [18]. Where that filtering happened matters more than the ratio. The 4,000 transformations were assembled from published literature [8], an agent narrowed them against chemical constraints [9], and the survivors were identified in computational simulations [10]. No physical synthesis or measured water uptake appears in the account [20]. Zheng credits the result to targeted design guided by model suggestions learned from literature-based community knowledge, rather than the brute-force linker screening the field usually pays for [11].
Set that against the size of the space. Zheng puts the number of possible MOF designs in the millions [12], which makes 4,000 documented transformations something under half a percent of even the low end of that estimate [19]. The search is not the scarce input. The written record of what has already been made, and how, is.
Which is why both papers land on the same unglamorous step. The models already hold the known rules of chemistry; what they lack, according to Cooper and Zheng, is the application of those rules through to simulating how a given molecule actually gets made [6]. Those instructions sit in journals, textbooks and footnotes accumulated over a century, and have to be collected, translated and fed in [7]. Cooper's own description of the target is a model that can read the instructions, mash them together, and say which combination is likelier to work [21]. The Matter paper he wrote with Kathryn Miller of NIST is the plumbing for that [15]: gather the constraints and practicalities, the polymers used, the dynamic bonds, the properties and applications discussed, tag it for machine reading, hand the filtering back to the machine, then find the flaws in execution and run it again [16].
Note the asymmetry between the two tracks. Zheng had a corpus before he had candidates [8]. Cooper's polymer effort is ongoing with the first step still being the collection of community knowledge [14]. Same institution, same premise, and one of the two is still upstream of any candidate list.
Zheng also points out that the autonomous lab has been an academic idea for decades, and that what makes his method possible now is large language models, trained rather like a graduate student handed a pile of instructions and allowed to learn from mistakes [13][17]. That is an honest account of which component changed. What changed is the reading, not the pipetting. A lab that procures the pipetting and inherits an undocumented archive of procedures has bought throughput it cannot feed, which is the entire content of calling data curation the first major step [5].
Ranked by verification strength, evidence, and original report placement.
Zhiling Zheng, an assistant professor of chemistry in Arts & Sciences at Washington University in St. Louis, proposed in a recent essay for the journal Science how AI systems can massively scale up the testing of model predictions.
Zheng: "We already see that AI is powerful in terms of predicting new structures."
Zheng wrote in the Science essay: "At the heart of this platform is the AI's ability to read chemistry like a chemist."
Christopher Cooper, an assistant professor of energy, environmental and chemical engineering at the WashU McKelvey School of Engineering, published a paper in the journal Matter documenting how to curate troves of data for polymer synthesis.
Cooper and Zheng describe data curation as the first major step of this work: data needs to be collected and converted to a form machine learning models can easily digest.
The models have been fed the known rules of chemistry, which is how they make predictions; what is missing is the application of those rules, following through to run simulations on making those molecules.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Single institutional source, simulation-only results
Everything rests on one phys.org write-up of two WashU researchers' own publications, with no independent reporting, no linked data or code, and no experimental validation. The central performance assertion - 10 materials beating state-of-the-art aluminum-based adsorbents - is a simulation output with no bench synthesis or measurement disclosed, and no hit-rate or failure statistics are given. The method description is specific and internally consistent, which keeps this above the floor, but nothing in the cluster is externally checkable.
Two papers and one curated library; no external users
Concrete artifacts exist - a Science essay, a Matter paper, the DPAL curated library run through case studies, and an agent pipeline that produced 10 simulated candidates - so adoption is not zero. But every artifact originates with the same two WashU groups, the polymer analogue is still at the data-collection stage, and no outside lab, vendor, or production deployment is named anywhere in the cluster.
Framing outruns the simulation-only result
The headline and framing - self-driving labs moving 'closer to materials discovery at scale', AI reading chemistry 'like a chemist', a methodology impossible before LLMs - sit well ahead of what is shown: ten computationally screened candidates from a 4,000-item literature corpus, with no synthesis, no measured performance, and a knowledge base covering a fraction of a percent of a design space the researchers themselves size in the millions. The overstatement is in scale and maturity rather than in fabricated detail; notably, the article's own framing of curation as the binding constraint is the understated part of the story.
Institutional research promotion, no adversarial voice
The piece is structured as research communication around two named faculty from one university, quoting only those faculty, with no external chemist, competing method, or skeptic included. Both researchers stand to benefit professionally from framing their own recent essay and paper as a step-change, and the outlet reproduces that framing without pushback. There is no disclosed commercial stake, funding source, or product to sell in the supplied material, which keeps this short of the top of the range.
Internally consistent but unreplicated and single-sourced
The narrative is coherent, the method steps are specific, and the two independent research threads corroborate each other on the central point that curation precedes prediction, which supports moderate confidence in what was reported. Confidence is held down by having one publisher, one institutional viewpoint, no primary documents in the cluster, no dates for the underlying publications, and no way to verify the comparative performance claim.
build
Fabricated SQLite CVEs cleared NVD, CISA ADP and Red Hat before anyone ran the code1 distinct publisher
build
Hybrid Post-Quantum TLS: Same Protocol, a 1,216-Byte Key Share1 distinct publisher
product
Illinois team builds the charge into the extraction molecule, and the reagent bill drops1 distinct publisher
security
NIST's multi-cloud tally: one resilience win against 23 new problems1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
phys.org
1 article · August 26, 2026