Skip to content

Science1 publisherNot yet confirmed elsewhere3 min readPublished

Researchers cite ten-company drug-discovery model to argue biomanufacturers should pool data

Roger Hart and Kelvin Lee cite a model ten drug makers trained on 2.6 billion pooled data points as the case for sharing data across biomanufacturing. Its gains were measured on drug-discovery assays, so the payoff on plant data is untested.

The Scientist · Science desk

How we use AISend a correction

Illustration accompanying Researchers cite ten-company drug-discovery model to argue biomanufacturers should pool data
Generated illustration

What happened

  • The ten consortium members were Amgen, Astellas, AstraZeneca, Bayer, Boehringer Ingelheim, GSK, Janssen, Merck KGaA, Novartis and Servier.
  • According to the paper's authors, the shared model significantly outperformed the models of every individual member.
  • Janssen's Wouter Heyndrickx reported a median Relative Improvement of Proximity to Perfection above 12% for conformal efficiency, with one result above 20%.

Compiled by The ScientistSomething wrong?How this is made

Why it matters

  • cost Before pooling can help a manufacturer, each member has to standardize its ontologies and schemas and connect sensors from several vendors, and it pays for that work itself.
  • constraint Regulators have not decided how shared models will be judged in submissions, so a manufacturer has no defined way yet to use one in support of a filing.
  • precedent Manufacturing consortia can copy a joint-training setup that rival drug makers have already accepted, so preparing the data becomes the harder hurdle.

The comparison behind the result fits the question a company asks before joining. Each member's own model is the control. According to Hart, Lee and colleagues, the pooled model significantly outperformed every one of them [1]. Heyndrickx's wording is narrower. He reported benefits in "most of the classification or regression tasks," and wrote that the model generally enhanced predictivity, sometimes substantially [2].

The 12% median is a relative measure, as its name says [3]. It does not mean twelve percentage points of accuracy. Turning it into hits found, or assays a chemist could skip, would take the baseline scores in the consortium's own papers.

The pooled data are also sparse. Treat each data point as one molecule-assay measurement. More than 21 million molecules crossed with more than 40,000 assays gives at least 840 billion possible pairs [12]. The 2.6 billion points fill at most about 0.3% of them [13]. That is an average of no more than about 124 measurements per molecule [14].

The thing this doesn't tell you is whether the gain carries over from molecules to a production plant. The MELLODDY model used "the chemical structure of potential drug compounds to predict how they will function" [7]. The groundwork the authors say manufacturing needs is a different job: common ontologies and schemas, and sensors from multiple vendors that communicate in real or near-real time [8]. Hart leads the Big Data Project at the National Institute for Innovation in Manufacturing Biopharmaceuticals [15]. He says redesigning legacy data infrastructure is expensive and time-consuming, and that equipment developers give multivendor operability low priority [9]. Regulators have not yet decided how such models will be evaluated in data submissions. People trained in both biomanufacturing and data science are scarce [10]. A drug-activity benchmark scores predictions on data that are already clean and pooled, and those barriers all come before that step.

Lee accepts that in-house work has its uses. "Many organizations are already pursuing in-house solutions," he told GEN. "The benefits are the ability to tailor the solution to one's specific situation." [11] His case for pooling is that firms have problems in common: "Many organizations face similar challenges and opportunities related to development and manufacturing," Lee said [16].

I think MELLODDY answers a narrower question well. Ten competitors put decades of proprietary data into one model trained on a secure platform [5][6], and the pooled model beat each firm's own [1]. To show the same for bioprocess data, someone would have to run that member-against-pool comparison on plant records. Until then, the team's warning that "Companies that do not adopt these capabilities are competitively disadvantaged from realizing the full benefits of their big data in the biopharmaceutical manufacturing marketplace" is an argument from analogy [17].

What to watch

  • A member-against-pool comparison run on bioprocess data, the test needed to show whether MELLODDY's design carries over to plants.
  • Regulatory guidance on how shared or federated models will be evaluated in data submissions.
  • Equipment vendors committing to multivendor sensor interoperability, which Hart says they currently rank low.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence40
Adoption20
Hype gap+35
Incentives50
Confidence40
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    The shared model's results significantly outperformed those of any of the individual members.

    ReportedSupportedSource: Hart, Lee and colleagues, as reported by GEN2 sources— create a free account to open themView cited source
  2. [2]

    Wouter Heyndrickx, machine learning scientist at consortium lead Janssen, reported benefits in "most of the classification or regression tasks," writing that the model generally enhanced predictivity, sometimes substantially.

    ReportedSupportedSource: Heyndrickx, separate article, as reported by GEN2 sources— create a free account to open themView cited source
  3. [3]

    The median Relative Improvement of Proximity to Perfection for conformal efficiency exceeded 12%, with one exceeding 20%.

    ReportedSupportedSource: Heyndrickx, as reported by GEN2 sources— create a free account to open themView cited source

Sources

1 independent publisher whose own reporting we read for this story.

  1. genengnews.com

    1 article · October 7, 2026

    Big Data Opportunities in Biomanufacturing

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Topics

Entities

Loading related stories