Skip to content

Science1 publisher3 min readPublished

The disclaimer quietly left the room: 40 million daily health questions now end with the user

A STAT opinion piece argues the measurable shift in medical AI is not benchmark wins over physicians but the removal of the safety text that used to route people to care.

The Scientist · Science desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • Several studies claim that AI outperforms physicians on clinical reasoning tasks.
  • More than 40 million Americans ask ChatGPT a health question every day.
  • A study published this year found that medical disclaimers, once standard in chatbot answers to health questions, have largely disappeared. The study is referenced in the opinion piece but not named there.
  • Today's leading models will not only respond to health questions but ask follow-ups and attempt a diagnosis.
  • Most ChatGPT health conversations happen outside clinic hours, and most lead not to a doctor but back to the user, someone with no training in what to provide, how to prompt, or how to critically appraise what comes back.

Compiled by The ScientistSomething wrong?How this is made

Why it matters

An opinion piece published by STAT on August 19 makes a claim that is easier to audit than any benchmark result: more than 40 million Americans ask ChatGPT a health question every day [2], and a study published this year found that the medical disclaimers once standard in chatbot answers to health questions have largely disappeared [3]. That combination matters because leading models now do more than answer: they ask follow-up questions and attempt a diagnosis [4], while most of those conversations lead not to a doctor but back to the user, someone with no training in what to provide, how to prompt, or how to critically appraise what comes back [5].

The framing the piece is arguing against is the recurring one: several studies claim AI outperforms physicians on clinical reasoning tasks [1]. Take the volume figure at face value and annualize it and you get roughly 14.6 billion health questions a year [1], which is a larger consultation surface than any health system on earth. The disclaimer finding is the part worth pressing on, and it is also the thinnest part of the source: the study is referenced but not named in the piece [3], so the effect size, the models tested, and the coding method are not available here. If it holds, it describes a design change with clinical consequences, made without notice, at that scale.

The piece traces where those answers land. Oura sells a 50-biomarker blood panel through Quest Diagnostics for $99 [6]. Function Health, valued at $2.5 billion in November, lets members order 160 lab tests a year, schedule a full-body MRI, and authorize ChatGPT to read the results [7]. Ro and Hims will write prescriptions for weight loss or anxiety after an asynchronous intake [8]. Doctronic, which calls itself the world's number one AI doctor, has run 24 million consultations and now writes AI-generated prescription refills in Utah [9]. Note the proportions: Doctronic's entire cumulative consultation history is smaller than a single day of ChatGPT health questions [2]. The regulated-looking corner of this market is the small corner.

On why medicine and not law or software, the author's answer is money: health care is close to one-fifth of the American economy [10]. That is a sufficient explanation for the speed, and it is also why the disclaimer change is unlikely to reverse on its own.

The author's technical objection is that clinical reasoning and model reasoning are not the same process, and that the hidden part of clinical reasoning, including abandoned hypotheses and the forks that an observer never sees, does not exist in the corpus these models train on [11]. That is an argument, not a measurement. The measurable adjacent fact is that mechanistic interpretability exists as a field because the engineers who built these models cannot fully explain them [12], and that interpretability remains nascent, particularly for clinical concepts and decisions [13].

Watch three things. Whether the disclaimer study is replicated with named models and dates, so the removal can be tracked release by release rather than asserted. Whether Utah's tolerance for AI-generated refills [9] becomes a template other states copy. And whether the licensing debate the piece describes, treating autonomous clinical AI more like a clinician than a device [14], produces any obligation to route users out of the chat.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories