Science1 publisherNot yet confirmed elsewhere3 min readPublished
Researchers cite ten-company drug-discovery model to argue biomanufacturers should pool data
Roger Hart and Kelvin Lee cite a model ten drug makers trained on 2.6 billion pooled data points as the case for sharing data across biomanufacturing. Its gains were measured on drug-discovery assays, so the payoff on plant data is untested.
The Scientist · Science desk

What happened
- The ten consortium members were Amgen, Astellas, AstraZeneca, Bayer, Boehringer Ingelheim, GSK, Janssen, Merck KGaA, Novartis and Servier.
- According to the paper's authors, the shared model significantly outperformed the models of every individual member.
- Janssen's Wouter Heyndrickx reported a median Relative Improvement of Proximity to Perfection above 12% for conformal efficiency, with one result above 20%.
Compiled by The ScientistSomething wrong?How this is made
Why it matters
- cost Before pooling can help a manufacturer, each member has to standardize its ontologies and schemas and connect sensors from several vendors, and it pays for that work itself.
- constraint Regulators have not decided how shared models will be judged in submissions, so a manufacturer has no defined way yet to use one in support of a filing.
- precedent Manufacturing consortia can copy a joint-training setup that rival drug makers have already accepted, so preparing the data becomes the harder hurdle.
The comparison behind the result fits the question a company asks before joining. Each member's own model is the control. According to Hart, Lee and colleagues, the pooled model significantly outperformed every one of them [1]. Heyndrickx's wording is narrower. He reported benefits in "most of the classification or regression tasks," and wrote that the model generally enhanced predictivity, sometimes substantially [2].
The 12% median is a relative measure, as its name says [3]. It does not mean twelve percentage points of accuracy. Turning it into hits found, or assays a chemist could skip, would take the baseline scores in the consortium's own papers.
The pooled data are also sparse. Treat each data point as one molecule-assay measurement. More than 21 million molecules crossed with more than 40,000 assays gives at least 840 billion possible pairs [12]. The 2.6 billion points fill at most about 0.3% of them [13]. That is an average of no more than about 124 measurements per molecule [14].
The thing this doesn't tell you is whether the gain carries over from molecules to a production plant. The MELLODDY model used "the chemical structure of potential drug compounds to predict how they will function" [7]. The groundwork the authors say manufacturing needs is a different job: common ontologies and schemas, and sensors from multiple vendors that communicate in real or near-real time [8]. Hart leads the Big Data Project at the National Institute for Innovation in Manufacturing Biopharmaceuticals [15]. He says redesigning legacy data infrastructure is expensive and time-consuming, and that equipment developers give multivendor operability low priority [9]. Regulators have not yet decided how such models will be evaluated in data submissions. People trained in both biomanufacturing and data science are scarce [10]. A drug-activity benchmark scores predictions on data that are already clean and pooled, and those barriers all come before that step.
Lee accepts that in-house work has its uses. "Many organizations are already pursuing in-house solutions," he told GEN. "The benefits are the ability to tailor the solution to one's specific situation." [11] His case for pooling is that firms have problems in common: "Many organizations face similar challenges and opportunities related to development and manufacturing," Lee said [16].
I think MELLODDY answers a narrower question well. Ten competitors put decades of proprietary data into one model trained on a secure platform [5][6], and the pooled model beat each firm's own [1]. To show the same for bioprocess data, someone would have to run that member-against-pool comparison on plant records. Until then, the team's warning that "Companies that do not adopt these capabilities are competitively disadvantaged from realizing the full benefits of their big data in the biopharmaceutical manufacturing marketplace" is an argument from analogy [17].
What to watch
- A member-against-pool comparison run on bioprocess data, the test needed to show whether MELLODDY's design carries over to plants.
- Regulatory guidance on how shared or federated models will be evaluated in data submissions.
- Equipment vendors committing to multivendor sensor interoperability, which Hart says they currently rank low.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence40
- Adoption20
- Hype gap+35
- Incentives50
- Confidence40
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
The shared model's results significantly outperformed those of any of the individual members.
ReportedSupportedSource: Hart, Lee and colleagues, as reported by GEN2 sources— create a free account to open themView cited source - [2]
Wouter Heyndrickx, machine learning scientist at consortium lead Janssen, reported benefits in "most of the classification or regression tasks," writing that the model generally enhanced predictivity, sometimes substantially.
ReportedSupportedSource: Heyndrickx, separate article, as reported by GEN2 sources— create a free account to open themView cited source - [3]
The median Relative Improvement of Proximity to Perfection for conformal efficiency exceeded 12%, with one exceeding 20%.
ReportedSupportedSource: Heyndrickx, as reported by GEN2 sources— create a free account to open themView cited source - [4]
The MELLODDY Consortium members were Amgen, Astellas, AstraZeneca, Bayer, Boehringer Ingelheim, GSK, Janssen, Merck KGaA, Novartis and Servier.
- [5]
The consortium pooled decades of proprietary data, some 2.6 billion data points spanning more than 21 million molecules and more than 40 thousand assays in on-target and secondary pharmacodynamics and pharmacokinetics.
- [6]
The consortium used a secure computing platform to jointly train a predictive drug-activity model on the pooled data.
- [7]
The shared quantitative structure-activity relationship model used "the chemical structure of potential drug compounds to predict how they will function."
- [8]
The researchers say preparation requires standardizing ontologies, schemas and data integration; secure shared datasets that protect proprietary data; interoperability and real-time or near-real-time connectivity among sensors from multiple vendors; and regulatory clarity, workforce development and shared business cases.
- [9]
Hart says redesigning legacy data infrastructure is expensive and time-consuming, and multivendor operability is a low priority for equipment developers.
- [10]
Regulators have not yet determined how such models will be evaluated in data submissions, and workers with combined expertise in biomanufacturing and data science are scarce.
- [11]
"Many organizations are already pursuing in-house solutions," Lee tells GEN. "The benefits are the ability to tailor the solution to one's specific situation."
- [12]
More than 21 million molecules crossed with more than 40,000 assays gives at least 840 billion possible molecule-assay pairs.
- [13]
Treating each data point as one molecule-assay measurement, 2.6 billion data points fill at most about 0.3% of the possible pairs.
- [14]
2.6 billion data points across more than 21 million molecules is an average of no more than about 124 data points per molecule.
- [15]
Roger Hart is senior fellow and Big Data Project lead at the National Institute for Innovation in Manufacturing Biopharmaceuticals; Kelvin Lee is a professor at the University of Delaware. They argue for standardizing and sharing key insights industry-wide without compromising proprietary information, based on MELLODDY results.
- [16]
"Many organizations face similar challenges and opportunities related to development and manufacturing," Lee says.
ReportedInsufficientSource: Kelvin Lee to GEN2 sources— create a free account to open themView cited source - [17]
"Companies that do not adopt these capabilities are competitively disadvantaged from realizing the full benefits of their big data in the biopharmaceutical manufacturing marketplace."
ReportedInsufficientSource: The research team, quoted by GEN2 sources— create a free account to open themView cited source
Sources
1 independent publisher whose own reporting we read for this story.
- genengnews.comBig Data Opportunities in Biomanufacturing
1 article · October 7, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.
Topics
Entities
- MELLODDY ConsortiumFollow
- Janssen PharmaceuticalsFollow
- NIIMBLFollow
- University of DelawareFollow
- Roger HartFollow
- Kelvin LeeFollow
- Wouter HeyndrickxFollow
- GEN (Genetic Engineering & Biotechnology News)Follow