Skip to content

Science1 publisher2 min readPublished

Nature urges regulators to test AI medical devices in real clinics before deployment

Nature's editors want AI medical devices tested in real clinics, noting that just 19 of some 4,600 clinical AI papers were prospective randomized trials. Until the FDA decides, hospitals need to check whether a vendor's accuracy figure came from real patients or from a simulation.

The Scientist · Science desk

Illustration accompanying Nature urges regulators to test AI medical devices in real clinics before deployment

What happened

  • The FDA published a discussion paper in August on how to regulate generative-AI-enabled medical devices and is taking comments until 19 October.
  • Under current risk-based rules, bandage makers can self-certify, stethoscopes can be lab-tested, and diagnostic products that affect clinical decisions face real-world testing.
  • Nature's editorial argues that AI scribes summarizing physician-patient conversations are more than administrative tools, because their outputs feed clinical decisions.
  • A September Comment in Nature Medicine argued that pre-registered clinical trials should become standard for AI in health, as they are for drugs and vaccines.

Compiled by The ScientistSomething wrong?How this is made

Why it matters

  • exposure A health system that treats regulatory clearance as proof of performance takes on the risk from any AI product that was certified without real-world assessment.
  • precedent If the FDA accepts that scribe outputs drive clinical decisions, scribes would move toward the risk tier where real-world testing, and sometimes trials, come before authorization.
  • decision Health systems and researchers holding their own deployment results can put them before the FDA while its approach is still at the discussion-paper stage.

The March review in Nature Medicine covered some 4,600 papers, and 23% of them used real-world patient data [2]. That comes to roughly 1,060 papers [1]. The 19 prospective randomized trials [3] are about 0.4% of the whole [2] and under 2% of the real-world subset [3]. I'd give the 19 more weight than the 23%. A study can use real patient records and still look backwards. A prospective randomized design compares patients whose clinicians used the tool with patients whose clinicians did not.

The thing this doesn't tell you is how many products now in clinics lack that kind of evidence. The review counts papers, not authorized devices. The editorial says some AI-powered medical products are being certified without real-world assessment [9], but it does not say how many, or report any case of patient harm. The figures on pace also need care. One estimate puts output at roughly three peer-reviewed clinical AI articles a day [12], and more than 230 million people a week ask ChatGPT health questions [13]. That second figure counts people using a general chatbot. The editorial treats it separately from the growing number of products built to assist clinicians, including with decisions [16].

Lower-risk devices can be lab-tested because they have a history to be judged against. A revised stethoscope or a new kind of bandage can be benchmarked against decades of data, including from real-world use, according to the editorial [10]. AI is powering devices that did not exist before, and clinical data for them are scarce or non-existent [11]. Companies often test these models in one-off simulated scenarios that measure the accuracy of the model's judgement [14]. A score like that belongs on a research leaderboard. A hospital has a different problem: whether a summary drawn from a real consultation, read by a busy clinician, changes an order or a diagnosis.

Asked whether generative-AI medical products need more comprehensive scrutiny, Nature's editors wrote: "The answer in most cases should be 'yes'." [15] They want the tools tested comprehensively and transparently before deployment [1]. I think health systems should apply that standard themselves without waiting for regulators. Before a scribe or diagnostic aid feeds decisions, a hospital should ask for a prospective evaluation on patients and workflows like its own. Until then, it should treat a vendor's simulation scores as screening data. The case rests on a gap in the published record. It would weaken if post-market data showed simulated accuracy closely tracking clinical performance.

What to watch

  • Whether the FDA's approach after the comment period puts AI scribes in the tier that requires real-world testing before authorization.
  • Whether any regulator makes pre-registered clinical trials a condition of authorization for clinical AI, as the September Nature Medicine Comment proposes.
  • A published count of authorized AI devices cleared without real-world data, to show whether the gap seen in papers holds for products.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories