Science1 distinct publisher3 min readPublished
More than 1,000 published studies rest on these methods, and new work makes their ambiguity a proven property of the mathematics rather than a suspicion about particular analyses. Where that leaves individual published claims is the open question.
The Scientist · Science desk

Compiled by The ScientistSomething wrong?How this is made
Lumpability dates to the 1960s: it specifies when the states of a Markov model can be grouped together without changing how the system behaves [7]. Tarasov and Uyeda arrived at it from the wrong end of the library. They were asking whether two shades of green apple should count as one colour or two, and ran into the same grouping problem while working on beetle anatomy [6]. What came back was more than a taxonomy rule. Every discrete-state Markov model, they show, has an equivalent hidden-state form built from simple identical components [8], and written that way, the model's symmetries sit on the surface where they can be analysed [9].
That is the step that mattered, because the trait-dependent diversification models had resisted direct analysis; their mathematics was too complex to interrogate [4]. The suspicion was not new. Earlier work had already shown that evolutionary models resting on entirely different assumptions about history can generate exactly the same observations [3], and practitioners had been watching these methods produce confident wrong answers for years without a cause to point at [2].
The stick insects are the part to keep. The original study asked whether male leg protrusions, used in fights over females, were associated with diversification, and found no evidence that they were; Tarasov and Uyeda agree with that reading [13][14]. When they fitted the same data across a broader set of models, standard statistical methods preferred scenarios in which the weapons did affect diversification [15]. The ambiguity, in this case, pushed toward the finding a journal would rather have [1].
The consequence is harder than a caveat about statistical noise. If two histories predict identical observations, no amount of that data separates them [2]. Model comparison scores fit, and fit alone; it has no way to flag a rival history that predicts precisely the same thing. That ceiling belongs to the method itself, and it is why the authors describe the misleading results as coming from ambiguity built into the models rather than from isolated mistakes in particular analyses [12].
The thing this work does not tell you is which of the more than 1,000 studies built on these models [1] would survive a wider model set, or how often the ambiguity bites hard enough to flip a sign. The paper locates the mechanism and marks the boundary of what current methods can say; it does not remove the ambiguity [16], and Tarasov is explicit that the whole problem is not solved [17]. Read the stick insect reanalysis as a calibration exercise rather than a verdict on anyone's beetles: a single dataset and a single trait showing that a compelling causal story can be the best-fitting story and still be wrong.
A second, quieter result follows. The decomposition is a statement about discrete-state Markov models in general, one of the most widely used classes of stochastic model in science [18], and Tarasov says the property had gone unnoticed through decades of work on them [11]. Phylogenetics is only where it surfaced first.</body_markdown> </invoke>
Ranked by verification strength, evidence, and original report placement.
Mathematical models that estimate how biological traits and environmental factors influence species formation and extinction have been used in more than 1,000 scientific studies.
These models carry a known weakness: they can sometimes lead scientists to wrong conclusions, and for years no one fully understood why.
Several years ago researchers found that many evolutionary models can generate exactly the same observations even when they rest on entirely different assumptions about evolutionary history.
Whether the same ambiguity affected the more sophisticated models used to study how traits shape biodiversity had remained unclear because their mathematics was too complex to analyse directly.
The study was carried out by Sergei Tarasov at the Finnish Museum of Natural History and Josef Uyeda at Virginia Tech, and is published in the journal Nature Communications.
The work began with a question about whether a red, a light green and a dark green apple should be grouped as two colours or three, a classification puzzle the researchers also met while studying beetle anatomy.
Distinct publishers with included, body-backed reporting in this cluster.
phys.org
1 article · September 2, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
science
Narwhal tusks hide two spirals twisting against each other, and the mismatch is the point3 distinct publishers
leadership
Dropping Item 407(j) left boards learning cyber risk from the people they supervise1 distinct publisher
science
Mount Sinai puts a youth protein on aging microglia, and the mice answer1 distinct publisher
science
Half the resistance genes in livestock manure also show up in 875 wild farm mice1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One release, one voice, one checkable paper
Everything on this page descends from a single institutional announcement that Phys.org reproduces, and every quotation in it belongs to Sergei Tarasov, one of the two authors. What keeps it above the level of a pitch is that the load it carries is a theorem in a peer-reviewed Nature Communications paper with a DOI a reader can pull, and the demonstration is specific enough to be argued with. What is absent is anyone outside the author list: no independent statistician, no phylogenetics group, and no direct word from the original stick insect team beyond Tarasov's report that they agree.
A thousand studies on the old methods, nothing yet on the fix
The number doing the rhetorical work — more than 1,000 studies built on these diversification models — is the museum's own count, undated and unaudited. On the other side of the ledger there is nothing: no software release, no reanalysis by an unaffiliated group, no journal or reviewer response. Adoption here points backwards at the exposed install base, not forwards at Hidden Expansion.
Careful text under a hard headline
The headline promises methods that 'give false answers'; the prose underneath is more disciplined than headlines usually allow — the authors endorse the original stick insect conclusion, concede the ambiguity is not removed, and say outright they have not solved the problem. The overshoot that remains is arithmetical: a 1,000-study figure sitting a few lines from a single reanalysis invites readers to conclude a large literature is wrong, which nothing here demonstrates.
The discoverer is also the narrator
A museum publicises its own researcher's paper, Phys.org passes the wording along largely intact, and the single quoted voice is a co-author explaining why his own theorem matters. No funder is disclosed, no product is being sold, and no rival lab is offered a line of reply — the pressure here is reputational rather than financial, and it runs in exactly one direction.
Firm on the mathematics, thin on the fallout
A claim about every discrete-state Markov model lives or dies on the paper, and that paper is peer-reviewed and citable, so we are reasonably steady there. The thing readers will actually want to know — which published diversification findings are now in doubt — rests on one worked example and one co-author's characterisation of what other researchers think. We would not lean any weight on that half.