Science1 distinct publisher3 min readUpdated
A model trained at NTNU did work its authors estimate would take a person 40,000 hours. The output: flowering times have moved 2.5 days a decade, and the tropics moved most.
The Scientist · Science desk

Compiled by The ScientistSomething wrong?How this is made
The Royal Botanic Gardens, Kew has published a State of the World's Plants and Fungi report that examines what digitization and machine learning are doing to botany, with contributions from 400 researchers in 40 countries [1][2]. The headline number inside it is an operational one: a classifier trained at the NTNU University Museum in Trondheim scored 8 million digitized herbarium specimens for whether the plant was in flower, a job its authors put at roughly 40,000 human hours [5][6][7].
That is about 20 working years of trained-eye labour, or a nominal 2,000 hours a year [8][11]. The machine took one week, according to postdoctoral researcher David Williamson [7]. Worth doing the division: 40,000 hours across 8 million sheets is 18 seconds per sheet [10]. The baseline is therefore a fast, narrow visual call on an already-photographed image, not curation, not identification, not label transcription. The saving is real, and it is a saving on one specific step.
The scale is also a fraction of what is sitting there. James Speed, a plant ecology professor at the same museum, says more than 145 million plant and fungi preparations have been digitized worldwide, across more than 170 institutions in 40 countries [3][4]. The 8 million sheets processed are about 5.5 percent of that [9], covering 200,000 species, an average of 40 sheets per species [6][12]. Williamson notes the technologies used are barely five years old, applied to collections built over centuries [24].
The result: global flowering times have shifted by an average of 2.5 days per decade over the past century, per Speed [13]. Over a century that is on the order of 25 days [23]. Flowering has moved both earlier and later, with the largest differences in the tropics [14]. Whether the 2.5 days is a net shift or an average magnitude of movement in both directions changes how you read it, and the material as supplied does not settle that [13][14]. The tropical finding surprised the team, because the largest warming has been in the far north [15]. Speed's explanation is that tropical flowering tracks precipitation, and a rainy season can arrive earlier, later, or not at all [16]. The consequence flagged in the report is a timing mismatch between flowering and pollinating insects, with knock-on effects for insect-eating birds and food production [17].
The same pipeline is aimed at the extinction accounting problem. More than 90 percent of fungal species remain unmapped [18], and only 1,000 plant species have been formally declared extinct, a figure the report treats as far too low [19]. Statistical models can estimate the probability that a species is extinct rather than merely undiscovered, and models can flag apparently unknown species in digitized collections for expert review [20].
Both quoted scientists put a ceiling on this. "Machines can't do the work of biologists, but they can help experts analyze datasets that would previously have taken a lifetime," Williamson said [21]. Kew taxonomist and report co-author Martin Cheek was blunter: "I think the potential for AI is enormous, but it is still currently potential" [22].
What to watch: whether the flowering classifier's error rate and validation set are published, since a phenology signal of 2.5 days per decade sits close to the plausible noise floor of a binary image call [13]; whether the remaining roughly 137 million digitized sheets get processed or stall on institutional access terms [4][9]; and whether any of the model-flagged candidate species reach a formal description, which still requires taxonomists [20][22].
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Speed says more than 145 million plant and fungi preparations have been digitized worldwide, from more than 170 institutions in 40 countries.
David Williamson, a postdoctoral researcher in machine learning for natural history at the NTNU University Museum, along with Speed and botanists from the Trondheim herbarium, trained a machine learning model to recognize whether plants are in flower on digitized herbarium specimens.
Williamson: "By way of comparison, the machine here took one week to do something that would have taken a human about 40,000 hours."
The source states that 40,000 hours equal about 20 working years.
Speed says the analyses showed that global flowering times have shifted by an average of 2.5 days per decade over the past century.
Statistical models can calculate the probability that a species is extinct as opposed to undiscovered, and AI can recognize unknown species in digitized collections and flag them for experts, speeding up the naming process.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Named researchers and an institutional report, but no methods, metrics, or second source
All figures trace to one publisher carrying an NTNU-sourced account, with named, affiliated researchers (Speed, Williamson) and a named Kew report co-author (Cheek), which is better than anonymous sourcing. Against that: no link or citation to the report or any peer-reviewed paper, no accuracy or validation metrics for the flowering classifier, no uncertainty range on the 2.5 day per decade figure, no treatment of herbarium sampling bias, and no independent corroboration anywhere in the cluster. The 40,000-hour human baseline is an explicit estimate ('about'), and dividing it by 8 million sheets yields an unexamined 18-second-per-sheet assumption. Credible attribution, thin verifiability.
One real large-scale run on top of a large but mostly inaccessible corpus
This is genuine deployment, not a demo: a trained model was actually run across 8 million specimen sheets covering 200,000 species and produced a published scientific finding, and it sits on a real global digitization base of 145 million-plus preparations from 170-plus institutions. Adoption is nonetheless narrow. The run covers roughly 5.5% of the reported digitized stock, less than 16% of the world's herbarium specimens are digitally accessible at all, the work is attributable to one museum group, and the wider extinction-triage and species-flagging uses are described only as capabilities with no deployed instance cited.
Speed-up framing outruns unpublished accuracy, but the source carries its own brakes
Mild overstatement. The framing ('AI optimists will be encouraged', '20 years of work completed in a single week') leans on an unaudited researcher estimate and never reports how often the classifier is right, which is the number that determines whether the 2.5 day per decade result holds. That pushes the gap positive. It is held close to zero because the same article publishes the deflating facts itself: Cheek's 'enormous, but it is still currently potential', the sub-16% digitization ceiling, gaps concentrated in biodiversity-rich countries, and Williamson's repeated insistence that models are only as good as their data and that experts are needed at every step.
Institutional promotion of the institutions' own work, partially self-disclosed
The material is an institutional science communication chain: NTNU University Museum researchers describing their own model and Kew describing its own flagship report, republished by an aggregator, with no adversarial or independent reporting in the cluster. Both institutions benefit from framing digitization and collections work as high-value and underfunded, and the piece closes by underscoring 'the critical importance of the work that humans are doing on museum collections'. No commercial vendor, funder, or product interest is disclosed or evident, and the sources volunteer their own limitations, which moderates the score rather than clearing it.
Plausible and internally consistent, but single-sourced and unverifiable as published
Confidence is capped by having exactly one publisher and one source document. The account is internally coherent, arithmetically consistent where checkable (40,000 hours to 20 years implies a 2,000-hour year), attributed to named experts at identifiable institutions, and unusually candid about its own limits, which supports moderate confidence in the descriptive facts. It is not enough to confirm the classifier's reliability, the derivation of the 2.5 day per decade figure, or how the 8 million sheets relate to the 145 million-specimen pool.
Distinct publishers with included, body-backed reporting in this cluster.