Science1 distinct publisher2 min readPublished
Computational biology proves a method by trying to break it on data it has not seen. Biomedical foundation models slip that test, and how the field responds decides what any benchmark number is worth.
The Scientist · Science desk

science
Rydberg chain spectra match Ising CFT predictions, turning a simulator into an instrument1 distinct publisher
science
The largest personality genome scan caps common-variant prediction at 13.3% of variance1 distinct publisher
science
Sleep tech turns rest into evidence, and employers are the ones who will need rules1 distinct publisher
science
BiteNetI annotates 14 ion types in an existing protein structure in seconds1 distinct publisher
Compiled by The ScientistSomething wrong?How this is made
Falsification works in this field because a claim and a test can be made to touch. A model that predicts an assay result asserts something checkable about cases it has not seen; when the cases come back the other way, the model is wrong, and the wrongness is the useful part. On the premise the piece sets out, that link loosens [4]: what is under test is a representation rather than a single prediction, and any benchmark reaches the representation only through some downstream task. Ten task scores tell you about ten tasks.
The reference list does hold the case where community assessment settled something. The entry for CASP round XIV is annotated as the independent evidence that AlphaFold2 reached near-experimental accuracy for many single-protein structure prediction targets [7]. Note the qualifiers: many, and single-protein. It worked because each target had an experimental answer that existed independently of the model. The DREAM assessment and the crowdsourcing tradition sit in the same list [10]. So do the systems now in question: AlphaFold 3, Boltz-2, AlphaGenome, Evo 2, a general-purpose pathology model, generative transformers for the natural history of human disease, and a roadmap for building a virtual cell [11]. For those, the abstract never says what would play the role of the experimental answer.
The accessible text is a statement of the problem rather than a resolution of it. It commits to discussing epistemological value, refutation versus verification versus utility, guiding principles, and the community's role [5], and it asks outright how the limitations of these models can be tested [6]. Cited alongside is Kapoor and Narayanan on leakage, the contamination that makes held-out scores look better than the model is [8].
On the community's role, the arithmetic is awkward. Reading the piece costs $39.95 as a single article, against the $21.58 per issue the same page quotes for a $259 annual subscription [12], roughly 1.85 times as much [13]. Four of the twenty references visible are preprints, two of which describe models in the debate [14].
My reading, and I will hold it conditionally: utility is a legitimate criterion, and a sufficient one for a deployment decision, but it is not evidence about mechanism, and scoring a model only on utility shows what it does, not that it has learned biology. The bar that earns the stronger claim is the CASP bar, which is targets whose answers are generated after the model is frozen and scored by people who did not train it. Whether its authors believe that bar is reachable for a virtual cell is left unanswered by the abstract [5].
Ranked by verification strength, evidence, and original report placement.
The reference list cites Stolovitzky, Monroe & Califano on the DREAM (Dialogue on Reverse-Engineering Assessment and Methods) pathway inference assessment, Ann. N.Y. Acad. Sci. 1115, 1-22 (2007), and Saez-Rodriguez et al. on crowdsourcing biomedical research.
An article titled 'Benchmarking biomedical foundation models' was published on nature.com; its abstract states that transparent evaluation and proof of reproducibility, generalization and replicability of algorithms are the bedrock of method development in computational biology.
The abstract states that many benchmarking efforts have been developed for problems ranging from structural biology to translational biomedicine.
The abstract states that rigor is relatively controllable for tasks such as the prediction of patient outcomes or the outcomes of biological assays, but that the problem is exacerbated when the aim is to benchmark foundation models.
The abstract states that the parameters constituting foundation models are supposed to capture the patterns underlying the data, and that the models are therefore parameterized embodiments of the phenomena that gave rise to the data.
The abstract says the piece discusses the epistemological value of foundation models; whether they can be refuted, verified or evaluated primarily on the basis of utility; what principles should guide their benchmarking; and what role the scientific community should play in that benchmarking process.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · September 3, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One page, abstract-deep
Everything checkable sits in the front matter of a single paywalled page: an abstract, a price list, and a bibliography that stops mid-entry. The abstract tells us which questions the authors take up and withholds their answers, so the benchmarking principles the piece advertises are unread here. What does hold weight is the reading list — CASP14, the sepsis external validation, the leakage paper are independently published work a reader can go and verify without Nature's permission.
Nothing to count yet
A position paper about how to grade models leaves no footprint we can measure. No consortium has adopted its principles, no benchmark has been rebuilt around them, and the page itself offers no citation or readership figures. If the field acts on this, the evidence will surface somewhere else, later.
A thesis built on front matter
The imbalance runs against our own reporting, not against Nature. The abstract poses questions — can these models be refuted, verified, or judged only by usefulness — and our framing answers them, declaring that biomedical foundation models slip the field's central test. Possibly right, but the sentences that would settle it are on the other side of $39.95. The gap stays modest because the one hard capability claim in view, AlphaFold2's near-experimental accuracy, is credited to CASP's independent assessors rather than to a developer's scorecard.
The journal grades its own bookshelf
Look at who is asking for scrutiny. The same journal family being urged to submit models to community assessment published AlphaFold 3, AlphaGenome, Evo 2, the pathology foundation model and the disease-trajectory transformers; the bibliography is substantially a Nature back catalogue. Then the reading fee: $39.95 for one article arguing that transparent evaluation is the bedrock of the field, close to double what a subscriber pays per issue. And the assessment lineage being recommended is also the most-cited one on the page — Stolovitzky's name appears on four visible entries, from DREAM to the self-assessment trap.
Single source, and it is the subject
Nobody else in our coverage has touched this, so there is no second account to catch a misreading — and one detail in our own reporting, the shape of the reference list, already disagrees with what the page actually shows. The individual citations are solid and traceable; the argument they support remains unread.