Skip to content

Build1 publisher3 min readPublished

VeriSim's injected patient noise costs seven open-weight models 15 to 25 accuracy points

VeriSim injects recall gaps, low health literacy and stigma-driven non-disclosure into simulated patients, and its authors report seven open-weight models shedding 15 to 25 diagnostic accuracy points while conversations run 34 to 55 percent longer.

The Engineer · Build desk

Illustration accompanying VeriSim's injected patient noise costs seven open-weight models 15 to 25 accuracy points

What happened

  • Conversations under the noise condition run 34 to 55 percent longer than they do on the clean cases.
  • The 7-8B models lose 1.4 times as much accuracy as the 70B-and-above class under the same injected noise.
  • Medical fine-tuning on standard corpora gave the models limited protection against patient communication noise.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • cost A turn-billed clinical deployment pays 1.34 to 1.55 times the conversation for a noisy patient, and it pays that multiplier while scoring below the clean benchmark it was bought on.
  • decision A team that qualified a 7B model on MedQA-style presentations has to requalify it on noisy input before that model talks to patients.
  • contradiction The two published abstracts describe truth preservation differently, strict adherence in one and substantial preservation in the other, and which one holds decides whether the drop is a noise-handling result at all.
  • constraint The abstracts leave out the model names and the clean baselines, so the 15 to 25 point band cannot be turned into an expected relative accuracy for any specific deployment.

The verifier works utterance by utterance. VeriSim extracts atomic claims from each candidate patient turn and judges them against a UMLS-grounded vector index built with BioLORD embeddings, and the judgement uses the retrieved atoms' structured clinical metadata, meaning drug class, anatomical site and treats-condition relations, not surface-text similarity alone [3]. That design targets the failure the paper attributes to prompt-driven simulators, which produce naturalistic responses but hallucinate symptoms and contradict the medical record [15].

The degradation figure is published in two different units. The arXiv listing says realistic noise reduces diagnostic accuracy by 15 to 25 percentage points [4]. The HTML version of the abstract says accuracy drops 15-25% [5]. Those coincide only if the clean baseline is 100. The abstracts report results across seven open-weight LLMs but list neither the models nor their clean scores, so a reader cannot convert one form into the other [21].

Conversation length rises 34 to 55 percent under noise [6]. For a turn-billed deployment that is 1.34 to 1.55 times the turns per encounter [18], at lower accuracy than the clean run. The paper's premise is that this is how real patients talk: they forget onset times, ramble about tangential topics, or withhold information because of stigma [16]. Its illustration is a myocardial infarction patient who says "my chest feels heavy, maybe since last week... or was it Tuesday?" instead of describing substernal pain radiating to the left arm over three days [22].

The size finding is the one that touches a procurement decision. The 7-8B models degrade 1.4 times more than the 70B-plus models [7], and the HTML abstract's phrasing, 40 percent greater degradation for 7B against 70B+, is the same ratio [8][20]. Apply that factor to the low end of the reported band and the small tier gives up roughly 21 points where the large tier gives up 15 [19]. For the number to transfer, your intake has to resemble these six dimensions and your flow has to tolerate the extra turns. The model you actually run also has to behave like one of the seven tested.

One of the six dimensions is stigma-driven non-disclosure [2]. Withholding removes a fact from the transcript. A longer context window or a better system prompt can recover a rambling patient's timeline. The same fixes do nothing for a symptom the patient chose not to mention. The reported band is a single range for the noise condition, so a team cannot yet tell how much of its own loss would come from silence and how much from disorder.

Medical fine-tuning on standard corpora gave limited robustness benefit against communication noise, according to the paper [9]. Those corpora are written in clinical register, which is the register the patient in the example conspicuously fails to use. Validation rested on two raters: a board-certified physician and a licensed nurse scored the conversations on truth, realism, clinical utility and noise fidelity, with inter-annotator agreement of at least 0.80 on every dimension, and an LLM judge closely tracked their ratings [10][11].

The HTML version calls VeriSim truth-preserving and says it maintains strict adherence to medical ground truth through hybrid UMLS-LLM verification [13]. The listing page says it substantially preserves each patient's medical record [12]. For an evaluation harness the distance between strict and substantial decides whether the drop measures a model handling garbled speech or a model handling a record that lost entries. The framework is released open source [17].

What to watch

  • Whether the released repo includes per-dimension ablations, separating loss from non-disclosure from loss from rambling.
  • A camera-ready that names the seven open-weight models and their clean baseline scores. A buyer needs both to map the band.
  • Whether fine-tuning on noisy transcripts closes the gap that fine-tuning on standard medical corpora left open.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories