Build1 publisherNot yet confirmed elsewhere3 min readPublished
A note checker with no accuracy figure, and the labelled dataset it borrowed to show its misses
Krasyn ran its clinical note checker against Omi Health's open benchmark and published the disagreements. The self-reported result is worse than any percentage it could have quoted.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- Krasyn declined to publish an accuracy figure for its clinical note checker, saying it has no clinician-adjudicated reference set to measure against.
- It instead tested the tool against Omi Health's MIT-licensed medical-note-eval, which ships synthetic primary-care dialogues, model-written SOAP notes and LLM-judge labels.
- The run covered 36 transcript and note pairs from six dialogues and six named frontier writers, through the production service with sentence IDs and citation tags stripped.
- On the ten notes Omi's panel had marked with a major unsupported claim, Krasyn's judge raised an Unsupported label in one.
- Twelve of the sixteen code flags came from a single rule that read a doctor's question as an undenied topic despite the patient answering it.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- precedent Counts and itemised disagreements against labels the vendor did not write is a standard a self-reported percentage cannot meet, and buyers in clinical documentation can now ask competitors for the...
- contradiction With LLM judges on both sides and no clinician adjudication on either, the disagreements do not establish which tool is wrong; they mark an unarbitrated boundary in what counts as supported by a...
- constraint The miss that matters is the plausible clinical conclusion, which the checker's own prompt policy admits on purpose, so closing it means redefining Supported rather than tightening a threshold.
- exposure Because a re-run can change what the tool says was omitted, the evidence sitting behind a signed note depends on which run produced it, and the vendor's answer is to store rather than recompute.
The two failures in this report do not behave alike, and that difference is the reason to read a disagreement table rather than a score.
The pertinent-negative rule is code with no model in it, and it failed identically on every note it touched. It located the doctor's question about blood clots, migraines with aura and uncontrolled blood pressure, and never registered the patient's "Correct, none of those" as the answer [11]. That class of error is cheap: enumerable and reproducible, so it can be fixed. Less comfortable is the precision underneath it. Only two flags in the whole run were correct, both on a weight change recorded as negative when it was asked about and never answered [12]. The arithmetic also does not close. Twelve flags are the denial bug, a thirteenth misread a phrase about foot swelling, two were right, which leaves one of the sixteen never described [19].
The other failure has no patch. In dialogue_2 the visit is history only, and two notes assert the cough is "likely related to allergic etiology"; all three of Omi's judges counted that a major unsupported claim and Krasyn's judge called it Supported [13]. The shape recurs three more times in the sample, including "Type 2 diabetes mellitus" written where the doctor said "your diabetes" or nothing at all [14]. Krasyn's explanation is that its judge prompt allows faithful clinical translation to count as Supported, and it adds the line that matters most: each of these reads as sensible, which is why the signing clinician would not catch them either [1]. Across all eighteen labelled notes, the tool applied the Unsupported label exactly once [18].
Against that, a headline percentage would have been a courtesy to the vendor. No clinician-adjudicated reference set exists to compute one from [3], so any figure would have been Krasyn grading itself, and the labels it borrowed instead are Omi's, published under MIT with the dialogues and the model-written notes [5].
The exercise is small and half-blind, and says so. Six synthetic dialogues, short and clean, with LLM judges on both sides and no clinician adjudication on either; the column that says which side looks right belongs to the vendor [16]. Only half the sample is comparable at all, because the 2026 benchmark ships per-writer totals rather than per-note counts [6], leaving eighteen of the thirty-six pairs with nothing to disagree with [20]. And the checker is not fully stable against itself. Verdicts held on seven of eight pairs run twice, while the omission list changed on four of eight, which is why that list is offered as a prompt to look rather than a count [15].
None of this makes the report weak. It makes it readable by someone deciding what to buy, because both error classes are visible separately instead of averaged into a single number that hides which one would reach their patients.
What to watch
- Whether the pertinent-negative rule ships a fix that removes the bad denial flags without suppressing the two that were correct.
- Whether anyone puts clinicians on the same 18 notes to adjudicate the allergic-etiology class of disagreement.
- Whether Omi restores per-note unsupported-claim counts for the 2026 benchmark, which would make the other 18 pairs comparable.