Skip to content

Science1 publisher3 min readPublished

Most quantum computing results do not re-run, so treat advantage claims as unverified

A Regensburg group assessed 4966 quantum computing papers and could only attempt replication on about a quarter. Of 127 checked by hand, roughly 11 yielded code that actually ran.

The Scientist · Science desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • Wolfgang Mauerer at the Technical University of Applied Sciences Regensburg in Germany and colleagues evaluated thousands of scientific papers on quantum computing and found that most currently cannot be reproduced; New Scientist reports this may undermine credibility of recent industry achievements and start a replication crisis in the field.
  • For quantum use cases to become routinely used they will have to work on many different quantum computers, so any demonstration of a truly valuable use of a quantum computer must be reproducible.
  • The analysis was done in two parts: first a manual evaluation of a curated sample of 127 papers from the past five years, then a computer program that automated and generalised the analysis to 4966 papers.
  • Five criteria were used: the first three focused on whether the paper included code that could be run on an independent quantum computer and how much instruction and documentation was provided; the fourth evaluated available hardware information; the fifth tested whether the provided quantum program could run without errors.
  • Among the 127 manually analysed papers, only 24.4 per cent provided code the team could even try running, and 64.5 per cent of those codes failed to execute successfully.

Compiled by The ScientistSomething wrong?How this is made

Why it matters

A group led by Wolfgang Mauerer at the Technical University of Applied Sciences Regensburg in Germany assessed thousands of quantum computing papers and found that most of the reported results cannot currently be reproduced, according to New Scientist [1]. That matters for anyone weighing a quantum pilot, because a useful demonstration has to work on many different quantum computers before it can be routinely used, which is precisely what reproducibility measures [2].

The work ran in two passes: a manual review of a curated sample of 127 papers from the past five years, then an automated version of the same assessment applied to 4966 papers [3]. Five criteria were used. Three covered whether the paper shipped code that could be run on an independent quantum computer and how much instruction and documentation came with it, a fourth covered the hardware information provided, and the last tested whether the program ran without errors [4].

The manual numbers are the ones to look at, because they include that execution step. Of the 127 papers, 24.4 per cent provided code the team could even attempt to run, and 64.5 per cent of that code failed to execute [5]. That leaves roughly 11 papers, about 8.7 per cent of the sample, that got as far as running at all [6]. The automated pass had no execution stage and found that 26.8 per cent of the 4966 papers carried enough information to attempt a replication [7], implying roughly 3635 papers, about 73.2 per cent, that did not [8].

This is not a first snapshot. Mauerer says the result is poor by the standards of conventional computer science, and that a smaller study his group did four or five years ago already found a bad situation that has not improved: "We thought the numbers would be better by now" [9].

The diagnosis changes how a buyer should read a failed replication. Mauerer attributes much of the gap to the machines rather than to authors, noting that conventional programmers can work abstractly without worrying whether their computer's physical characteristics change from one day to the next, while today's quantum machines, including cloud-accessible ones, are unusually variable [10]. On that reading, a paper that will not reproduce is not necessarily careless; it may be a result that held for one device on one day. The operational consequence is the same either way.

Two researchers quoted by New Scientist read the finding as normal for the field's age. Fred Chong at the University of Chicago says he is not particularly alarmed, that reproducibility becomes more of a priority as a field matures, that conventional computer science took decades to insist on such standards, and that innovation may matter more than reproducible infrastructure for quantum computing at this time [11]. William Zeng at the Unitary Foundation says he is not surprised, and expects communities of maintainers plus agentic coding, which he says already makes it easier to generate reproduction code directly from a paper's result, to improve matters [12].

That context is worth holding alongside the underlying uncertainty about value. Only a few problems are known for certain to need a quantum computer, and what else these machines are good for is an open question, with groups exploring everything from molecular simulation to airline logistics optimisation [13].

What to watch: whether the reproducibility rate moves at all, given that it did not budge over the previous four to five years [9], and whether the stated appetite for change converts into venue requirements for code, hardware detail and execution artefacts. Team member Ralf Ramsauer says the response to the study has so far been positive [14]. For procurement, the near-term rule is unglamorous: a published advantage claim is a hypothesis about your hardware until you have re-run it on your hardware.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories