Science1 distinct publisher3 min readPublished
Leeds researchers confirmed their predictions with pea and potato proteins bought off the shelf, which is the right first check and also the easiest one to pass, leaving most of the list untested and the formulation work still ahead.
The Scientist · Science desk

Compiled by The ScientistSomething wrong?How this is made
The severity of the filter deserves a number. Read "tens of millions" at its lowest plausible value, 10 million, and nearly 800 retained sequences is 0.008 percent of the input, about one in 12,500 [1]. A larger starting pool only makes that fraction smaller. A filter that aggressive is useful in proportion to how right it is about what it keeps and, much harder to establish, about what it throws away.
On the first half, the Leeds group reports agreement: several commercially available proteins were tested, and proteins from peas and potatoes performed as the model predicted [10]. That is the correct first experiment, because those proteins can be ordered and put on a rig within weeks. It is also the least demanding version of the test. The account does not give the number of proteins tested, and it does not report what happened when proteins the model ranked as unlikely emulsifiers went through the same measurement [17]. Without that second arm, the confirmation establishes that the ranking does not contradict the known behaviour of ingredients already in commerce. It does not yet bound how often the model is wrong about the candidates on the list that, as Anwesha Sarkar puts it, had never previously been considered for the purpose [9].
The design is the part worth studying. Rik Sarkar of Edinburgh describes emulsifiers as often carrying a chemical structure called a diblock, which the team modelled mathematically for plant proteins and turned into machine-learning features drawn from statistical physics simulations [8]. The simulation asks how a protein attaches at the boundary between oil and water [5]. The learned features fingerprint the particular segments that govern that attachment [6] rather than scoring whole proteins, and that substitution is what makes a pool of tens of millions tractable at all [1].
Simha Sridharan's own framing is that candidates were never scarce, that testing them was the expense, and that no reliable predictor existed for which plant proteins behave like animal ones [11]. The screen speaks directly to that gap. What it leaves untouched is the distance between a ranked sequence and an ingredient: extraction at usable purity and a defensible price for crops that no one currently processes for protein. Phys.org's framing is that the work will cut years of costly trial and error [14]; that saving is plausible at the discovery stage and unmeasured past it, since no emulsion performance figures such as droplet size or storage stability appear in the account [18]. As triage for interfacial behaviour the pipeline earns attention from anyone assembling a plant-protein portfolio; as a claim about time to product it currently rests on two crops that were already on the shelf [10].
Ranked by verification strength, evidence, and original report placement.
Sridharan and Anwesha Sarkar used a simulation model to understand how proteins attach between oil and water mixtures, which is crucial for them to act as emulsifiers.
In collaboration with Rik Sarkar, the team applied machine learning to fingerprint specific segments of the protein that dictate their attachment behaviour.
By combining machine learning and statistical physics the team screened plant proteins to identify those that would resemble the emulsification performance of animal proteins in a fraction of the time needed for conventional experiments.
Rik Sarkar said emulsifiers often have a characteristic chemical structure called diblocks, that the team modelled this structure mathematically for plant proteins, and that using machine-learning-based features obtained from statistical physics simulations they can predict which plant proteins are most likely to work as natural emulsifiers.
Scientists from the University of Leeds' School of Food Science and Nutrition developed an AI process that can rapidly identify plant proteins capable of acting as emulsifiers from tens of millions of initial candidates.
To date the process has identified nearly 800 promising plant proteins.
Distinct publishers with included, body-backed reporting in this cluster.
phys.org
1 article · September 3, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
invest
Singular Photonics raises $2.15M, and the sales line matters more than the round1 distinct publisher
science
A 104,196-membrane screen moves the carbon-capture bottleneck to the lab bench1 distinct publisher
science
Petermann's 30-square-mile calving was the small one: two bigger sections are already cracked2 distinct publishers
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Peer-reviewed paper, single-channel telling
The load carried here is thin in an unusual way: the underlying work is peer reviewed and citable by DOI, which is more than most announcement-driven stories offer, yet every number a reader sees - tens of millions in, nearly 800 out, 'several' proteins tested - reaches us through one Phys.org write-up of the university's own framing. The experimental check is asserted qualitatively and no measurement accompanies it, so the strongest evidence in the story is a citation rather than a result.
Nothing past the bench
There is no uptake to measure. A paper exists and two supermarket-grade proteins behaved as the model said they would; no company, product, licence or trial appears anywhere in the reporting, and the researchers' own statement is that firms 'could be' interested. Scoring adoption from that would be inventing it.
Headline outruns the bench
Two sentences do most of the overreach. 'AI identifies nearly 800 promising plant proteins' converts model rankings into findings, and the flat prediction that the discovery will cut years of costly research is stated in the publisher's voice with no timeline behind it. Set that against the arithmetic - the screen keeps something like one sequence in 12,500 - and the gap is not that the work is weak but that a shortlist is being described as a result. The researchers' own quotes are noticeably more careful than the framing around them.
University comms with an alt-protein centre attached
This is a research-announcement story travelling the announcement channel, and the promotional pull is visible without being hidden: the supervising author co-directs the National Alternative Protein Innovation Center based at the same university, the report volunteers that food and cosmetics companies could be interested, and the sustainability framing lands before any measurement does. Funding and licensing arrangements go unmentioned, which is the one thing a reader would need to size the interest properly.
Trustworthy about what it says, silent on the rest
Confidence splits cleanly. That the method exists, was published, and matched predictions on pea and potato is well attested and easy to check against the DOI. How well it generalises, how often it is wrong, and whether any of the 798 remaining candidates matter are questions this reporting does not touch - and with a single publisher relaying a single announcement, there is no second account to close the gap.