Science1 distinct publisher3 min readPublished
Almost every tool in a 50-program audit at Barcelona's CRG reads a quiet stretch of DNA as an important one, so its harm scores partly track how often a site mutates rather than what mutating it does.
The Scientist · Science desk

Compiled by The ScientistSomething wrong?How this is made
A conservation score is an inference drawn from an absence, and that gap is where the models can go wrong. Nearly every tool in the audit asks one question of a position in the genome: has this stayed the same over millions of years? If it has, the software concludes the site was kept that way for a reason and that changing it must be bad [5]. Two unrelated situations produce the same stillness. Selection may have removed every variant that appeared there, or almost nothing may ever have appeared, because the site mutates rarely to begin with [9]. A model that cannot separate those carries an estimate of mutability inside its estimate of harm [6].
The heterogeneity involved is large and already partly mapped. Weghorn's earlier work put the mutation rate at the starting points of genes 35% above other regions [8], and other groups have turned up hotspots elsewhere [9].
The design decision worth underlining is the control. Ben Lehner's group, also at the CRG, made its mutations physically, hundreds of thousands of them across 500 human protein fragments, and measured the functional damage each one caused [12]. That readout comes entirely from direct measurement, independent of evolutionary conservation, which is what makes it usable as a check. It confirmed a genuine effect pointing the same way as the models' bias [13], and first author Hossameldin Ali reads that as genomes having evolved robustness against their own most frequent errors [20]. The Weghorn group then corrected for the real effect, and the skew was still there [15], which is what separates a claim about biology from a claim about software.
Scale is what lets the paper speak about gene families rather than curiosities: 13.5 million mutations across 6,659 genes [1] averages roughly 2,030 variants per gene [1]. On that base, the study reports DNA repair, cilia and sperm-function genes as likely overcalled for danger, and genes tied to intellectual disability and to conditions inherited from a single parent as likely undercalled [10]. Weghorn's summary of the two directions is "Both would be bad" [11].
How far this moves any real report is the open question, and the paper stops short of settling it: it can name the direction of the bias but not its size in any single result, and the score is one input among several that clinical experts use [17], alongside its other role in guiding treatment options and laboratory work on mechanism [19]. Weghorn's own expectation is that some variants could flip from harmless to harmful once mutation-rate biology is included, while most would shift more gradually [16]. My read is that this is a calibration fault rather than a collapse of validity, and for a specific reason: the error has a known direction and a control that measured it. What the audit does not hand you is a per-tool ranking or a per-case error rate, so nobody looking at a mid-range score today can say how far off it sits.
Ranked by verification strength, evidence, and original report placement.
A research team led by Dr. Donate Weghorn at the Center for Genomic Regulation (CRG) in Barcelona tested 50 of the world's top variant prediction tools against 13.5 million mutations spread across 6,659 human genes.
The work was published in the American Journal of Human Genetics.
Almost all the programs tested overestimate the potential harm of mutations that are less likely to occur than average, and underestimate the effect of more likely mutations.
The tools tested include AlphaMissense, built by Google DeepMind, along with EVE and popEVE, co-developed by other research groups at the CRG.
Almost all the computer programs look at a region of DNA and ask whether that spot has stayed the same over millions of years of evolution; if it has, the software decides it has been conserved for a reason, must matter, and that changing it must be bad.
The study finds the models are biased because some regions in the human genome can be much more mutation-prone than other regions.
Distinct publishers with included, body-backed reporting in this cluster.
phys.org
1 article · September 3, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
build
Anthropic's protein run is checkable, which is rarer than the 26.8% hit rate2 distinct publishers
invest
Meta FAIR says the standard way to plan a training run costs 10x more than it needs to1 distinct publisher
product
Google's new Flash buys its benchmark wins with extra tokens1 distinct publisher
invest
Washington pitches Carolina Principles to G20, urging no new AI rules or bodies1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Peer-reviewed and physically benchmarked, but single-channel
The spine of this is strong for a science story: a named journal paper with a DOI, an audit spanning 13.5 million mutations, and a yardstick that is measured rather than predicted — Lehner's lab built the mutations and recorded the damage. The researchers also tested the obvious alternative explanation and reported that the skew outlived it. What holds the number down is that all of it reaches us through one phys.org write-up tracking the institute's own account, with the dissenting developers paraphrased, no independent re-run, and not a single per-tool figure to inspect.
Audited tools are everywhere; the correction is nowhere
Two different things are being counted and only one of them exists in the world. The predictors under examination are in live use — described as the field's top 50, named down to AlphaMissense, EVE and popEVE, with clinicians reading their scores toward treatment choices and labs using them to chase mechanisms. The mutation-rate-aware rebuild the authors prescribe has no release, no patch and no adopter in this reporting; it is a recommendation in a discussion section. The audit itself is the only concrete event.
The brakes are on in the same paragraph as the finding
A claim that almost every leading predictor in clinical genomics is systematically miscalibrated could have been written far louder. Instead phys.org keeps the authors' own limits in frame — no inference to any individual patient, software as one input among several, mostly gradual shifts rather than reversals — and gives the sceptical developer view space before the reader reaches the fix. The framing sits a little under what the audit's scale would license, which is the rarer failure mode.
Auditor, benchmark and remedy share one address
Follow the institutional wiring. Two of the audited models, EVE and popEVE, were co-developed at the CRG; the measured data the audit is graded against comes from Lehner's lab at the CRG; and the prescribed fix runs straight through Weghorn's own prior result that gene start points mutate 35% more often. That does not make the finding wrong, and there is something creditable about a group testing its neighbours' tools alongside DeepMind's — but the story has essentially one address, and phys.org's copy stays inside the institute's framing rather than sourcing around it.
Direction firm, magnitude unresolved
I would bet on the direction: conservation-based scores partly track how mutable a site is, the effect is uneven across gene classes, and the honest robustness signal is too small to explain it. I would not yet bet on what that does to a report leaving a diagnostic lab. The size question is precisely where the reporting thins out — unnamed developers say modest, the authors say gradual for most variants, no per-tool numbers are shown, and no second outlet or independent group has weighed in.