Skip to content

Science1 publisher3 min readPublished

Forty-four percent of surveyed scientists report their bottleneck has moved downstream to the bench

A new report combines a survey of 637 scientists with 15 million Gemini conversations. Its authors trace the gap between mathematics and biology to what it costs to check whether an answer is right.

The Scientist · Science desk

Illustration accompanying Forty-four percent of surveyed scientists report their bottleneck has moved downstream to the bench

What happened

  • A report from Google, Google DeepMind and MIT FutureTech combines a survey of 637 scientists with analyses of 15 million Gemini conversations and an inventory of more than 2,600 specialized models.
  • About 44 percent of those scientists said their main research bottleneck had shifted downstream over the past two years, toward later stages such as physical experimentation and data collection.
  • Forty-one percent said their backlog of untested hypotheses had grown.
  • AI companies have spent recent months publicizing machine-generated proofs, some of which drew accusations of scooping academics, as evidence of scientific capability.

Compiled by The ScientistSomething wrong?How this is made

Why it matters

  • constraint Speedups land where an answer can be verified cheaply, so time saved generating hypotheses in biology queues behind lab capacity instead of turning into results.
  • contradiction Zou treats the jam as missing automation hardware and Acemoglu treats it as missing ground truth; a buyer can fix the first with equipment and cannot fix the second at all.
  • cost Clearing a bench bottleneck is a capital expense, and the public bet so far is $20 million to one university lab.
  • decision Funders and agencies writing AI acceleration into research plans now have to name which stage of the pipeline they expect to get faster.

The cost of checking is where the report's authors put the asymmetry. Mihai Codreanu, a research economist at Google and a lead author, said "89% of those who save time using AI spend more than a 10th of the time that they're saving actually checking AI outputs," and added that "46% spend more than a quarter." [8][6][7] For that second group, ten hours of saved work returns at least two and a half hours of verification, leaving seven and a half at best. [27] Codreanu said the checking burden shapes which tasks AI can take and how useful it is field by field: "There is a lot of heterogeneity." [10][9] A mathematical proof can be checked against formal rules. A predicted protein function has to be built in a lab and tested in a living system. [11]

The survey findings are self-reports, and they are modest in absolute terms: 44 percent of 637 is about 280 scientists, and 41 percent is about 261. [25][26] Respondents came from online panels, so frequent AI users may have been overrepresented. [3] The Scientific American account does not break the 637 down by discipline, and field-level splits are the substance of the heterogeneity claim. [28] A shift in where a scientist reports their bottleneck is also not a count of experiments run. [4]

The researchers quoted point to different causes for the jam. [29] James Zou, a Stanford biomedical data scientist who has built fully virtual automated labs and handed research to tools including DeepMind's AlphaFold, described a hardware limit. [12] "They're still fairly limited in the kinds of experiments that can be automated, mostly some chemistry experiments," Zou said. [13] "But things that involve, let's say animals? It becomes much harder." [14] He also said existing automated labs can be prohibitively expensive. [15] Daron Acemoglu, the Nobel laureate MIT economist who studies AI's labor impacts, located the problem in the questions themselves. "When you look at the tasks where AI has made much progress, they are the ones where you have ground truth," he said. "In most of medicine, there is no ground truth." [17][18] The first diagnosis is a capital problem. The second says some questions have no checkable answer to automate against.

Some of the slowness is deliberate. Clinical trials move at the pace of regulatory review, and any experiment involving people is constrained by safety checks before, during and after. [19]

Money is moving toward the bench anyway. Julius B. Lucks, principal investigator of the DREAM Cloud Lab at Northwestern University, is trying to automate more of the protein-engineering cycle, and his lab received $20 million from the National Science Foundation over the summer as part of a new national network of programmable cloud laboratories. [20][21] The same agency told researchers in a memo last week that AI will "transform" the questions they can answer and how science is done. [22]

What to watch

  • Whether the report's survey is repeated with a sampling frame that is not an online panel, and with results broken out by discipline.
  • Whether Northwestern's NSF-funded cloud lab publishes per-experiment cost and cycle time, the numbers that would show whether automation relieves the bench constraint.
  • Whether any follow-up measures published output or compounds tested, instead of where scientists say their bottleneck sits.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories