Science1 distinct publisher3 min readPublished
A team adapting the implicit association test to reasoning traces found four of five models working harder on association-incompatible prompts, which puts a measurable bias signal in the process rather than only in the answer.
The Scientist · Science desk

Compiled by The ScientistSomething wrong?How this is made
The human version of this test runs on a stopwatch, and the inferential trick is that timing is hard to curate: what you measure is the cost of sorting a pairing rather than the sorting decision itself [11]. Reasoning models offer something close to that, because they emit a variable quantity of intermediate text before committing to an answer [7]. Counting those tokens gives a meter reading on the same prompt an output audit would score once [1][3].
That is the design's contribution, and also where the caution belongs. The authors report convergent validity: RM-IAT effects predicted biases in a word-association task and a decision-making task already known to surface LLM bias [6]. Predicted, not produced. Nothing in the abstract shows the extra tokens causing the skewed answer. The two measures move together, which is what you want from an instrument and not the same thing as a mechanism.
The Claude 3.7 Sonnet result is the most useful part of the paper for anyone planning to use this. One model of the five reversed the pattern, 20% of the sample [5][1][2], and the reversal does not arrive with a story about that model being cleaner. A thematic analysis of its traces tied the flip to an unusual internal preoccupation with reasoning about bias and stereotypes [5]. So a token gap on association-incompatible prompts can register friction or it can register deliberation, and the score by itself will not distinguish them.
The thing the abstract does not tell you is size. Direction and consistency are stated; effect sizes, prompt counts and variance estimates are not in the public text [10], and the full paper sits behind a $39.95 purchase or a $119 annual subscription [9]. What is free is the apparatus: the analysis code on GitHub, an archived version on Zenodo, and a capsule bundling code, data and environment on Code Ocean [8]. The reproduction path costs less than the reading.
Wherever reasoning tokens are the billed unit, this stops being only an ethics result. A systematic asymmetry in reasoning length means the same task consumes more compute when its content cuts against a learned association, and that overhead tracks what users type rather than which model was procured. The abstract puts no number on the gap [10], which is exactly the number a budget owner would ask for first.
Conditioned view: I would run the RM-IAT as a screen on any reasoning model headed into a triage or screening workflow, since prior bias work has concentrated on outputs [2] and this looks at the route instead. I would also read a sample of traces before believing the sign of the score, because in this study one model in five earned its number the opposite way [2].
Ranked by verification strength, evidence, and original report placement.
A paper published in a Nature Portfolio journal introduces the reasoning-model implicit association test (RM-IAT) to study bias-like processing differences in reasoning models.
Previous research on bias in large language models has focused primarily on outputs.
The RM-IAT measures reasoning-token counts as an index of computational effort, capturing processing efficiency differences analogous to response latency differences in the human implicit association test.
Across o3-mini, DeepSeek-R1, gpt-oss-20b and Qwen3-8B, the authors find consistent evidence that association-incompatible tasks require greater computational effort than association-compatible tasks.
Claude 3.7 Sonnet exhibited reversed patterns, which a thematic analysis linked to its unique internal focus on reasoning about bias and stereotypes.
The authors report convergent validity with model outputs: RM-IAT effects predicted biases in two tasks known to capture LLM biases, in word association and in decision-making.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 31, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
science
Ammonia from untreated seawater and air at room temperature, with the numbers attached1 distinct publisher
science
The self-driving lab is out. Whether AI shows up in your filing is still open.1 distinct publisher
science
Nature Perspective: autonomous agents need internal states they must keep in range1 distinct publisher
build
The judge went synthetic first, which tells you which part of your pipeline is next1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Peer-reviewed and reproducible, but only the abstract is readable
This clears a bar most bias findings do not: a Nature Portfolio journal, code on GitHub, a Zenodo archive of the exact version and a Code Ocean capsule with the data and environment inside it. Then the paywall closes. What a reader can see reports that incompatible prompts took more effort, consistently, across four models — and not one effect size, prompt count or variance estimate. The reproducibility is genuine; the verification has simply been deferred to whoever pays.
No use beyond its authors is on record
Publishing code makes uptake possible; it is not uptake. Nothing in this reporting shows a replication, an independent audit, a lab adopting the RM-IAT internally or any product wired to it, and we will not convert the authors' own release into an adoption figure.
Disciplined in the paper, tidier in the retelling
Nature's wording is careful — 'bias-like processing differences', not bias — and our own framing holds that line. The overreach sits in how much weight the analogy carries: reasoning tokens standing in for human response latency is asserted in a sentence and defended somewhere behind the paywall. And the single model that went the other way is accounted for after the fact by a thematic reading of its own reasoning, which is the shape an explanation takes when it was not predicted. Modestly ahead of what anyone can currently check.
The commercial interest belongs to the publisher, not to a model vendor
The plainest stake here is the price list Nature prints under the abstract: $39.95 for a single copy, $119.00 a year for the journal, $32.99 per 30 days for Nature+. Beyond that, a team that has just named an instrument has the ordinary academic interest in seeing it catch on, which is also why they shipped the capsule. What is absent is the usual distortion — no lab is funding this, selling against it or being sold by it, and the five models named include ones from competing vendors.
One abstract, no second read
Every assertion we make traces to a single publisher's abstract. The code that would settle the numbers exists and nobody in this reporting has run it; no other outlet has examined the method; the anomaly and its explanation come from the same authors. That is where confidence should sit for an un-replicated paper whose body we have not seen — enough to report the direction, not enough to lean on it.