Skip to content

Invest1 publisher3 min readPublished

Kenya's LLM trial hangs on roughly 196 treatment failures among 9,347 patients

In the Penda Health trial, LLM support improved documentation and treatment plans while 14-day failure ran 2.2% against 2%, and the FDA's discussion paper on how to evaluate such tools is still only collecting comments.

The Investor · Invest desk

Photograph accompanying Kenya's LLM trial hangs on roughly 196 treatment failures among 9,347 patients
Photo: nature.com

What happened

  • A Nature Medicine trial published on 26 June had 103 clinical officers at 16 Penda Health facilities in Nairobi and Kiambu counties treat patients with and without large language model decision support.
  • The AI arm produced better documentation and more appropriate diagnoses and treatment plans, process gains that did not show up in the patient outcome.
  • In the 2024 RAPIDx AI trial, the six-month composite of cardiovascular death, myocardial infarction or unplanned readmission was 26% with AI and 26.4% with standard care across 3,029 patients.

Compiled by The InvestorSomething wrong?How this is made

Why it matters

  • constraint The largest trial in this material could not resolve a 23% odds reduction sitting on about 196 events. To prove an outcome benefit at a 2% failure rate, a study has to be far larger than 9,347 patients.
  • decision A hospital can put documentation quality and procedure volume into a model today; the patient-outcome case has no number to enter. Workflow economics is what the buying decision runs on.
  • contradiction Cryptopolitan casts the FDA as scrutinising the proof gap, while the agency says its own paper indicates nothing about future regulatory expectations. A vendor shipping next quarter meets no new evidence bar.
  • capability The consistency threshold splits a diagnostic queue: 49.4% of cases clear automatically at 98.9% accuracy, and the other 50.6% still needs a clinician. That caps the staffing saving at roughly half the caseload.

The raw rates and the adjusted estimate in the Penda Health trial point in opposite directions. Failure at 14 days was two tenths of a percentage point higher in the arm with LLM support [1], and the adjusted odds ratio still came back below one, in AI's favour [5]. Adjustment is carrying the direction of that estimate. At a pooled 2.1% failure rate, about 196 of the 9,347 patients in the primary analysis failed treatment by day 14 [2], and 344 of the 9,691 enrolled, 3.5% of them, never reached that analysis [3].

The adjusted estimate implies 23% lower odds [4]. Spread across roughly two hundred events, that is a few dozen patients, and the trial reported the difference as not statistically significant [5]. I'd call that underpowered on the outcome, not a demonstration that LLM support does nothing. The opposite reading sits in the same study: documentation and the appropriateness of diagnoses and treatment plans improved, and the patient outcome did not [7]. No serious adverse events were associated with the intervention [6].

For a countable saving in this material, go back to the older trial. RAPIDx AI's six-month composite differed by 0.4 of a percentage point across 3,029 patients [8], about 12 patients [5], while AI-supported care sent non-type-1 MI patients for invasive coronary angiography 47% less often [9]. A procedure count is something a finance office can price, once it has the base rate and a unit cost.

A vendor will quote the vignette gains. In a multi-country randomized trial under Nicholas Rounding, GPT-4o lifted the clinical vignette performance of 249 doctors by 18% in Kenya, 10.7% in Indonesia and 7.2% in the Netherlands, at P<0.001 [10]. The cases were simulated, the control group could not use the internet or clinical protocols, and harms were not evaluated [11]. The Kenyan uplift is 2.5 times the Dutch one [6]. A spread that wide depends on what the control clinician was otherwise allowed to consult.

Li Zhang, Jakob Nikolas Kather and colleagues reported 90.04% on seven diseases in Nature Medicine [12], the highest accuracy figure here, and it comes off a benchmark. The authors said the approach needs confirmation through trials in real clinical settings [14].

The FDA's 18 August discussion paper covers risk assessment, premarket assessment and postmarket monitoring, and it takes public input until 19 October [15], a window of 62 days [8]. The regulator has not set a bar: the agency says the document is not a draft guideline, not a proposed policy change and not an indication of future regulatory expectations [16]. An earlier FDA inquiry into real-world performance raised model drift and the limits of static metrics [17]. Cryptopolitan, citing the Financial Times, quoted the headline "Medical AI has a proof problem" [18].

That account frames the gap as a problem for hospitals buying these systems, for companies trying to show usefulness, and for authorities deciding what evidence counts [19]. It names no hospital that declined a purchase and no payer that withheld payment on outcome grounds. On the evidence here, the purchase a buyer can defend is the one justified by documentation quality and procedure volume. I'd revise that if a trial sized for a 2% event rate reported a significant reduction in 14-day treatment failure.

What to watch

  • What the FDA does after comments close on 19 October, and whether postmarket monitoring picks up outcome endpoints rather than static accuracy metrics.
  • Whether any vendor publishes a cost endpoint alongside process endpoints, the way RAPIDx AI's angiography frequency did.
  • Whether the Zhang and Kather agent holds 49.4% retention at 98.9% accuracy outside the seven-disease benchmark, in a live clinic.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories