Science1 distinct publisher2 min readUpdated
In a Harvard, MIT and University of Washington experiment, 228 reviewers deferred to an LLM at the same rate with or without a rationale. Only the silent version improved their accuracy.
The Scientist · Science desk
Compiled by The ScientistSomething wrong?How this is made
The rationale bought no additional compliance. Deference ran at about the same level whether the model argued its case or simply stamped a verdict, so what the written reasoning changed was the composition of the errors, not the amount of following [18]. That is the result with operational consequences, and the researchers say it contradicts the common assumption that model explanations augment human decision-making [15].
The 21-point gap between how often these reviewers went along with machine advice and how often they went along with human advice is the figure worth carrying into any screening design [17]. Direction matters more than volume, though. A rejection commits no budget and creates no owner, and the researchers reach for Xerox shutting down early Ethernet and PostScript work as their illustration of that cost, against Google Glass and Amazon's Fire Phone on the false-positive side [9]. Only one of those two failure modes leaves a project to write a post-mortem about.
The mechanism the authors offer is about accountability, not model quality. People weigh negative information more heavily than positive, a pattern the paper labels negativity bias [10], and as they put it, "Rejection is an active, eliminative decision that feels more consequential and accountable than preserving optionality," while also maintaining the status quo and requiring no resource commitment [11]. A fluent paragraph explaining why a proposal falls short hands the reviewer a ready-made justification for the decision they were already disposed toward, and the study reports evaluators leaning on surface cues such as fluency and coherence [12]. The authors describe narrative explanations as suppressing productive overrides, meaning the cases where someone independently verifies a persuasive model output before agreeing with it [16][4].
That distinction is where enterprise exposure sits. Screening under time pressure with only basic information is precisely the job LLM recommenders are being handed [19]. A stored rationale for each decision is an attractive artifact because it looks like a control and files like evidence. What it evidences is that a reason was available, not that anyone tested it.
The study's ground truth deserves stating plainly: correct means agreeing with four human experts on one challenge's submissions, so this measures alignment with a small panel rather than commercial outcomes [3]. Within that frame, the authors end on design rather than warning, writing that "effective human-AI collaboration requires designs that preserve rather than supplant independent human judgment" [13]. On this evidence, the explanation layer is a persuasion feature that has been filed under governance.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Researchers associated with Harvard Business School, MIT and the University of Washington ran an experiment in which 228 experienced evaluators assessed nearly 50 submissions to an MIT challenge.
The experiment tested three scenarios: human-only proposals with no AI assistance; LLM evaluations with a written rationale for the decision; and black-box AI pass/fail recommendations with no accompanying explanation.
Evaluators' decisions were compared to decisions made by four human experts, which were treated as the correct baseline.
Responses were classified as outright compliance with the LLM recommendation, overrides, or productive overrides, meaning the evaluator independently verified persuasive model outputs before deciding.
Overall, the evaluators accepted LLM recommendations 67% of the time.
Evaluators agreed with both black-box and narrative LLM decisions roughly 75% of the time, but agreed with human decisions only 54% of the time.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Single-source account of one unlinked study
The cluster contains exactly one article, which reports specific, internally consistent quantitative results (67% acceptance, ~75% versus 54% agreement, three conditions, 228 evaluators, four-expert baseline) and direct researcher quotes. But there is no link, venue, or peer-review status for the paper, no model identification, no effect sizes or per-condition breakdowns, and no independent commentary or replication, so the evidence base is one journalist's summary of one experiment.
No adoption data supplied
The supplied source contains no releases, deployments, benchmarks, usage disclosures, or named adopters. Its only adoption-adjacent statement is a general assertion that time-constrained decision-makers are 'increasingly turning to LLMs', with no figures or examples, and the experiment itself is a research setting rather than a production deployment.
Modestly overstated beyond one experiment
The headline framing ('Explaining itself made the AI more persuasive and the reviewers less right') is broadly faithful to the reported numbers, and the article does preserve nuance about task type and error cost. Overstatement comes from scope: a single innovation-screening experiment, scored against four experts as ground truth, is generalized into enterprise-wide guidance about high-stakes AI-assisted decisions, and the counterintuitive 'black-box is better' result is presented without uncertainty ranges or replication.
Academic authors, trade-press amplification
The findings originate with researchers at three universities rather than a vendor, and the conclusions cut against commercial explainability marketing rather than promoting a product, which limits promotional incentive. Offsetting factors are structural rather than disclosed: the reporting outlet serves an enterprise IT audience with an appetite for counterintuitive AI-risk findings, no funding or conflict disclosures appear, and no vendor or independent voices are quoted for balance.
Plausible and specific, but unreplicated and unlinked
Confidence is moderate-low. The mechanism is coherent with well-known behavioral effects and the numbers are specific and internally consistent, so the direction of the finding is credible. But everything rests on one publisher's account of one unlinked experiment, with an unverified expert baseline, no model identification, and no adoption evidence, and the load-bearing 'black-box beats narrative' result has no independent corroboration in the supplied material.
product
Cinemas, classrooms and ICE: smart glasses now need a venue-policy contingency1 distinct publisher
product
MIT's magnet trick makes correlated microwave signals without the cryostat1 distinct publisher
science
Flow control gets a shared benchmark, and a 38% friction cut nobody had to simulate first2 distinct publishers
product
MIT's bacterial transistors do one calculation every eight hours. That is the useful part.1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 22, 2026