Science1 distinct publisher3 min readUpdated
A Regensburg group assessed 4966 quantum computing papers and could only attempt replication on about a quarter. Of 127 checked by hand, roughly 11 yielded code that actually ran.
The Scientist · Science desk
Compiled by The ScientistSomething wrong?How this is made
A group led by Wolfgang Mauerer at the Technical University of Applied Sciences Regensburg in Germany assessed thousands of quantum computing papers and found that most of the reported results cannot currently be reproduced, according to New Scientist [1]. That matters for anyone weighing a quantum pilot, because a useful demonstration has to work on many different quantum computers before it can be routinely used, which is precisely what reproducibility measures [2].
The work ran in two passes: a manual review of a curated sample of 127 papers from the past five years, then an automated version of the same assessment applied to 4966 papers [3]. Five criteria were used. Three covered whether the paper shipped code that could be run on an independent quantum computer and how much instruction and documentation came with it, a fourth covered the hardware information provided, and the last tested whether the program ran without errors [4].
The manual numbers are the ones to look at, because they include that execution step. Of the 127 papers, 24.4 per cent provided code the team could even attempt to run, and 64.5 per cent of that code failed to execute [5]. That leaves roughly 11 papers, about 8.7 per cent of the sample, that got as far as running at all [6]. The automated pass had no execution stage and found that 26.8 per cent of the 4966 papers carried enough information to attempt a replication [7], implying roughly 3635 papers, about 73.2 per cent, that did not [8].
This is not a first snapshot. Mauerer says the result is poor by the standards of conventional computer science, and that a smaller study his group did four or five years ago already found a bad situation that has not improved: "We thought the numbers would be better by now" [9].
The diagnosis changes how a buyer should read a failed replication. Mauerer attributes much of the gap to the machines rather than to authors, noting that conventional programmers can work abstractly without worrying whether their computer's physical characteristics change from one day to the next, while today's quantum machines, including cloud-accessible ones, are unusually variable [10]. On that reading, a paper that will not reproduce is not necessarily careless; it may be a result that held for one device on one day. The operational consequence is the same either way.
Two researchers quoted by New Scientist read the finding as normal for the field's age. Fred Chong at the University of Chicago says he is not particularly alarmed, that reproducibility becomes more of a priority as a field matures, that conventional computer science took decades to insist on such standards, and that innovation may matter more than reproducible infrastructure for quantum computing at this time [11]. William Zeng at the Unitary Foundation says he is not surprised, and expects communities of maintainers plus agentic coding, which he says already makes it easier to generate reproduction code directly from a paper's result, to improve matters [12].
That context is worth holding alongside the underlying uncertainty about value. Only a few problems are known for certain to need a quantum computer, and what else these machines are good for is an open question, with groups exploring everything from molecular simulation to airline logistics optimisation [13].
What to watch: whether the reproducibility rate moves at all, given that it did not budge over the previous four to five years [9], and whether the stated appetite for change converts into venue requirements for code, hardware detail and execution artefacts. Team member Ralf Ramsauer says the response to the study has so far been positive [14]. For procurement, the near-term rule is unglamorous: a published advantage claim is a hypothesis about your hardware until you have re-run it on your hardware.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
For quantum use cases to become routinely used they will have to work on many different quantum computers, so any demonstration of a truly valuable use of a quantum computer must be reproducible.
The analysis was done in two parts: first a manual evaluation of a curated sample of 127 papers from the past five years, then a computer program that automated and generalised the analysis to 4966 papers.
Five criteria were used: the first three focused on whether the paper included code that could be run on an independent quantum computer and how much instruction and documentation was provided; the fourth evaluated available hardware information; the fifth tested whether the provided quantum program could run without errors.
Among the 127 manually analysed papers, only 24.4 per cent provided code the team could even try running, and 64.5 per cent of those codes failed to execute successfully.
The larger automated analysis did not include an execution step and found that only 26.8 per cent of the almost 5000 papers provided enough information to attempt a replication.
Mauerer says the result is rather negative compared to standards for traditional computer science studies, that a smaller-scale study four or five years ago already found the situation was bad and it did not improve over time, and that 'We thought the numbers would be better by now'.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Quantified method, single outlet, underlying study not linked
The reporting is unusually specific for a single-source item: a stated two-part design, corpus sizes of 127 and 4966, five named criteria, and three separate percentages, plus two independent outside experts who engage with the result rather than merely react. It is capped by the absence of the study itself: no title, venue, DOI, preprint link or peer-review status is given, the paper-selection method is described only as a curated sample, no per-criterion or per-venue breakdown appears, and no second publisher corroborates any figure.
Reproducibility practice rare across ~5000 papers
Adoption here is measured as uptake of reproducible-artifact practice in the quantum literature, and the two audits directly quantify it: 26.8 per cent of 4966 papers cleared the information bar for an attempted replication, and in the hand-checked sample only about 8.7 per cent shipped code that actually ran. Mauerer adds that a smaller study four to five years earlier found the same and that the numbers did not improve, so the low level is not a snapshot artifact. Uptake of the study's own remedy (its reproducibility-package template) is asserted by a co-author but not evidenced.
Counts solid, crisis-and-credibility framing runs ahead of scope
The numbers themselves are reported plainly and hedged with 'may', so this is not a large gap. The overstatement is one of scope: the audit scores academic paper artifacts on code, documentation and hardware metadata, yet the framing extends to undermining credibility of industry achievements and, in the cluster headline, to treating advantage claims as unverified, without a single vendor claim being re-run. Mauerer's own hardware-variability explanation and Chong's field-immaturity counterpoint both argue that low artifact reproducibility is partly a property of immature machines rather than evidence that reported results are wrong.
Author-promoted audit, insider commentators, no disclosures
Both people describing the finding are its authors, and the article closes with them pointing readers to their own reproducibility-package template, which is an interest in the problem being seen as urgent. The two counterweight voices are field insiders whose framings also favour their positions: an academic architecture researcher arguing innovation should outrank reproducible infrastructure, and a quantum non-profit figure whose remedy is maintainer communities and AI coding agents. No funding, grant or commercial affiliation is disclosed for anyone quoted, and the outlet's own framing choice is the attention-maximising one. Nothing in the material shows a paid or vendor-directed message, which keeps this mid-range.
Internally consistent but unverified and single-publisher
Confidence is moderate. The quantitative core is coherent, arithmetically consistent, disclosed as to method limits, and stress-tested by two named experts within the same piece. It is held down by structural thinness: one publisher, no access to or identification of the underlying study, an unspecified sampling method for the 127-paper curation, no execution testing in the large arm, and no independent corroboration of either the percentages or the claimed reception.
science
The self-driving lab is out. Whether AI shows up in your filing is still open.1 distinct publisher
science
Selenium's metabolite zoo finally gets one naming system, two decades late1 distinct publisher
science
AI weather models cannot forecast what they never saw. A hybrid method aims at that gap.1 distinct publisher
invest
Pleasant, and lonelier: a 12,365-person trial cuts against the AI companion pitch1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 17, 2026