Science1 publisher3 min readPublished
Researchers publishing in Nature intervened directly on the confidence representation inside language models and watched abstention move the other way, making a refusal rate something with a control input, at least on multiple-choice facts.
The Scientist · Science desk
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Clean causal design, thin reporting
The manipulation is what carries this: confidence was raised and lowered inside the network and abstention moved the other way, with mediation analysis behind the mechanism claim, which is more than a correlational study of calibration could deliver. Everything we hold, though, is the abstract and the first pages, twice from the same journal. No model is named, the only quantitative statement about the confidence term is a ratio, and the task is four-option multiple choice from a factuality set.
No deployment signal
Nothing here touches use. There is no release, no product setting refusal thresholds this way, and no operator reporting a changed abstention rate. The work stops at the bench, and scoring uptake would mean inventing it.
Vocabulary outruns the stimulus
Our own framing, that a refusal rate now has a control input, is what the steering result actually supports, and the paper hedges its bigger reading as 'consistent with' metacognitive control rather than proof of it. The stretch is in the language: 'metacognition' and 'autonomous agents' are heavy freight for behaviour measured on four-option factual questions, and the effect size a sceptic would check is given only as a comparison.
Disclosures not in view
Authorship, funding and competing-interest statements all sit in parts of the paper we do not have. The one interest visible in the text is the novelty claim staked in the introduction, which is too little to score against.
One document, seen twice
Peer review at Nature counts for something, and the design answers the question it poses more directly than the calibration literature it distances itself from. Against that, every claim we carry traces to a single paper delivered to us in duplicate, and the details that would let a reader argue with it, model coverage and effect magnitudes, are outside the text.
Compiled by The ScientistSomething wrong?How this is made
Activation steering is where the causal weight sits. Correlational work cannot separate the model gating on its own confidence from confidence and abstention both tracking how hard the question is; an intervention can. The authors pushed the internal confidence signal up and down, and abstention moved inversely, falling when confidence was boosted and rising when it was suppressed [3]. Their mediation analysis is the second half of that argument, indicating the behavioural change travelled through a redistribution of confidence rather than around it [3].
Phase 2 supplies the shape of the policy. Models behave as though they hold an implicit threshold on internal confidence, and the confidence term outweighs alternative mechanisms by roughly an order of magnitude [2], which puts those alternatives near a tenth of the effect size [11]. The abstract states that ratio without absolute values [12], so the direction of control is better established here than its gain. Earlier work in this area got abstention by bolting a threshold on after the fact or by fine-tuning for it [9]; the claim being tested here is that the native signal was already doing the job [10].
The awkward part for anyone planning to tune this is which signals are actually reachable. Verbal confidence, elicited in a separate forward pass, independently predicted abstention in every model tested, even though it discriminated correct from incorrect answers less well than the token probabilities did [5]. The abstract leaves those models unnamed [16]. Activation decoding indicated that both measures are lossy read-outs of a richer internal representation [6]. So the dial within reach is a proxy for the quantity doing the work, and part of the gate runs on the poorer indicator of being right. Push a refusal threshold up through that component and you remove answers; it does not follow that you remove the wrong ones.
This result is scoped to four-option multiple choice drawn from a factuality dataset [7]. A fixed option set with one correct answer gives confidence something clean to be about, and gives the experimenters a correctness label to score discrimination against. An agent deciding whether to call a tool, or a model deciding whether to answer a clinical question of the kind the paper cites as its motivating case [14], has neither. The authors themselves frame the work as mattering because models are becoming autonomous agents that must recognise their own uncertainty [8], a statement about why the question is worth asking, separate from evidence that the multiple-choice result carries into that setting.
On what is supplied: abstention on these items is a threshold policy over an internal confidence representation [15], the representation can be moved on purpose [3], and instructing a model to abstain at a stated confidence level shifts its behaviour in the instructed direction [4]. That is enough to treat refusal rate as a quantity with a control input. Specifying a target rate and holding it inside an agent loop is a different question, one this design did not test.
Ranked by verification strength, evidence, and original report placement.
Researchers developed a four-phase paradigm to test whether large language models use confidence signals to decide whether to answer or abstain; Phase 1 elicited baseline confidence without offering an abstention option.
In Phase 1, the models completed four-option multiple-choice questions drawn from a factuality dataset.
The abstract and opening of the paper refer to results holding 'across all models' without naming which models were tested or how many.
Phase 2 showed that large language models apply an implicit threshold to internal confidence when abstaining, with confidence effect sizes roughly an order of magnitude larger than alternative mechanisms.
Phase 3 provided causal evidence via activation steering: boosting confidence decreased abstention and suppressing confidence increased it, with mediation analysis confirming confidence redistribution as the primary mechanism.
Phase 4 instructed models to abstain at different confidence levels and found they adjusted their behaviour accordingly, indicating they read out and act on confidence to set abstention policies.
science
Rydberg chain spectra match Ising CFT predictions, turning a simulator into an instrument1 publisher
science
The largest personality genome scan caps common-variant prediction at 13.3% of variance1 publisher
science
Sleep tech turns rest into evidence, and employers are the ones who will need rules1 publisher
science
A neck bypass for Alzheimer's reached hundreds of hospitals on the strength of one video1 publisher
Publishers with included, body-backed reporting in this cluster.
2 articles · September 6, 2026