Science1 distinct publisher3 min readPublished
Sixteen people from a dozen fields spent three days on why research uncertainty gets lost on the way to the reader. Their most portable output is a definition with two halves, only one of which usually survives.
The Scientist · Science desk
Compiled by The ScientistSomething wrong?How this is made
Start with the evidence class, because it governs how much weight the conclusions can bear. CIFAR brought together a room of practitioners for a conversation, not a controlled study [2]. Sixteen people were in it [16], and what came out is agreement among practitioners plus worked examples from their own fields. There's no control wording and no comprehension measure, and the workshop does not include a before-and-after trust score [18]. That doesn't invalidate the conclusions so much as mark them as hypothesis-generating, with a central hypothesis that happens to be unusually testable.
Here it is, in the group's own working definition: uncertainty is everything not yet known about the thing you are trying to describe, and how much that missing knowledge could change the answer [7]. Read that as two clauses rather than one sentence. The first clause is the hedge, and hedges are cheap to write. The second is a magnitude claim, and it is the clause a reader needs in order to do anything [17]. A finding that names what is missing without bounding how much it matters has satisfied the etiquette of caution and none of its function.
The wildfire example is the cleanest demonstration of why that is structural rather than verbal. A wildfire scientist in the room argued that calling fires simply unpredictable risks abdicating responsibility: nobody can say when a given fire starts, how fast it moves or what it takes, but it is near-certain that fires will happen somewhere, and treating the whole thing as unknowable becomes an excuse not to plan [9]. The word "unpredictable" is doing the damage by collapsing a specific, bounded ignorance into a general one. Meanwhile the planning obligation does not move; local authorities still have to build mitigation on the assumption that a fire is coming [10].
The BICEP2 case cuts the other way, and it is the more disciplining of the two. In 2014 a team announced it had glimpsed the universe's first instant, and the signal turned out to be consistent with dust in our own galaxy [14]. Physics demands extraordinary confidence before it will say discovery, which the announcement had. What confidence about a number does not cover is whether you are measuring the right thing at all [14]. That's a category error, not a statistical one, and no error bar can capture it.
Which is why the statistician's point in that room lands harder than it sounds: the line between significant and nonsignificant, conventionally drawn with a p-value at 0.05, is as much a social convention as a mathematical one [13]. The thing a threshold crossing does not tell you is how big the effect is, or whether the measurement was pointed at the quantity you care about.
So the reader's tests fall out of the group's own categories: whether the account states a magnitude of doubt or only its existence, whether it distinguishes "we cannot say when and where" from "we cannot say anything," and whether it says whether the result was modeled in the lab or observed in the world, which the group flags as its own species of uncertainty, along with not knowing how people will respond to the information [11]. Genetic counselors do this translation as their job, moving population-level risk into something an individual can decide with [12], and an investigative journalist made the same distinction from the other end: beyond a reasonable doubt is not absolute certainty, and wrongful convictions happen when a court forgets the difference [15].
My view, with its condition attached: I would trust a result that states how wrong it could be over one that clears 0.05 quietly. The condition is that this comes from experienced communicators reasoning together [1], not from any measurement of what readers actually understood.
Ranked by verification strength, evidence, and original report placement.
Four authors spent three days in a room with a dozen others, including a cosmologist, an emergency management scholar, a genetic counselor, a hydrologist, a neuropsychologist, a health risk communication scholar, a cognitive neuroscientist, a theater director, an investigative journalist and a meteorologist.
The group convened earlier this summer through CIFAR, a research organization that brings interdisciplinary groups together on hard questions, for a workshop on how to grapple with and communicate about uncertainty in research.
The four authors come from four different disciplines: a sociologist, a human rights scholar, an Indigenous community health scholar and a microbial geneticist.
Each participant brought an object capturing their own struggle with representing uncertainty, including a handheld weather monitor, a blank flip chart, and a gachapon capsule toy bought without knowing which toy you will get.
The authors write that the instinct for many researchers is to treat uncertainty as a failing, something to paper over before anyone notices, while their conversation suggested the opposite: uncertainty is everywhere in every field and is worth naming out loud rather than hiding.
Most participants agreed that their struggle to articulate uncertainty makes it harder to share their work clearly with the public, and that glossing over it can damage public trust when people later discover the doubt that was left out.
Distinct publishers with included, body-backed reporting in this cluster.
Follow any of these and your For You feed starts watching them — no settings page required.
science
HIPAA Covers Less Than You Think, And "Anonymized" Is Not A Legal Shield1 distinct publisher
science
Whale-watch footage catches 1,417 humpbacks feeding along a corridor theory calls a fast1 distinct publisher
science
Police AI reached the evidence file before anyone measured how often it fabricates1 distinct publisher
science
Claude's watermark is a compliance artefact, not a cheating detector1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
First-hand, and only hand
The strength here is proximity: the people making the claims were the people in the room, and they say plainly what was said and by whom. The weakness is that nothing else in the piece does any lifting. Participant agreement stands in for a finding, the field anecdotes illustrate rather than test, and the one external study — qualitative research from March 2026 — arrives with no author, journal, or link a reader could chase. Even the checkable references, BICEP2's galactic dust and the 0.05 threshold, are true things borrowed as analogies, not support for the workshop's conclusion.
No trace outside the room
We have no basis for a number. The two-part definition is reported as the output of one workshop; nobody is shown using it — no journal guidance, press office, agency style rule, or newsroom practice is cited as having taken it up, and the piece does not claim any. Whether this framing travels is exactly the question the reporting leaves open.
Restrained, with one causal reach
This piece declines the move it could easily have made. It does not announce a fix; it says navigating uncertainty is the job and never fully goes away, which is a smaller and more honest promise than most workshop write-ups make. The overreach is narrow: the claim that glossing over doubt damages public trust when people later find it is presented as something most participants agreed on, and agreement is not a measurement. The definition itself is undersold rather than oversold — it is the most portable thing in the text and gets a single sentence.
Participants marking their own homework
Four attendees write the public account of the workshop they attended, on a platform built for academics to publish their own work, and the convener gets a warm mention with no scrutiny. That is a real self-assessment loop and it explains why no dissent from the room reaches the page. It scores low rather than high because there is nothing to sell: no product, no funding round, no policy ask, no claim that would move a budget.
Sure what was said, unsure it works
Two different levels of certainty are in play. What happened in that room — the disciplines, the objects, the definition, the wildfire argument — we can report with little hesitation, because the people who said it wrote it down. Whether the advice changes how uncertainty survives a handoff is untested and single-sourced, and one account of one three-day conversation leaves no second version to check the framing against.