Build1 distinct publisher3 min readUpdated
Krasyn ran its clinical note checker against Omi Health's open benchmark and published the disagreements. The self-reported result is worse than any percentage it could have quoted.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
The two failures in this report do not behave alike, and that difference is the reason to read a disagreement table rather than a score.
The pertinent-negative rule is code with no model in it, and it failed identically on every note it touched. It located the doctor's question about blood clots, migraines with aura and uncontrolled blood pressure, and never registered the patient's "Correct, none of those" as the answer [10]. That class of error is cheap: enumerable and reproducible, so it can be fixed. Less comfortable is the precision underneath it. Only two flags in the whole run were correct, both on a weight change recorded as negative when it was asked about and never answered [11]. The arithmetic also does not close. Twelve flags are the denial bug, a thirteenth misread a phrase about foot swelling, two were right, which leaves one of the sixteen never described [18].
The other failure has no patch. In dialogue_2 the visit is history only, and two notes assert the cough is "likely related to allergic etiology"; all three of Omi's judges counted that a major unsupported claim and Krasyn's judge called it Supported [12]. The shape recurs three more times in the sample, including "Type 2 diabetes mellitus" written where the doctor said "your diabetes" or nothing at all [13]. Krasyn's explanation is that its judge prompt allows faithful clinical translation to count as Supported, and it adds the line that matters most: each of these reads as sensible, which is why the signing clinician would not catch them either [14]. Across all eighteen labelled notes, the tool applied the Unsupported label exactly once [19].
Against that, a headline percentage would have been a courtesy to the vendor. No clinician-adjudicated reference set exists to compute one from [2], so any figure would have been Krasyn grading itself, and the labels it borrowed instead are Omi's, published under MIT with the dialogues and the model-written notes [4].
The exercise is small and half-blind, and says so. Six synthetic dialogues, short and clean, with LLM judges on both sides and no clinician adjudication on either; the column that says which side looks right belongs to the vendor [16]. Only half the sample is comparable at all, because the 2026 benchmark ships per-writer totals rather than per-note counts [5], leaving eighteen of the thirty-six pairs with nothing to disagree with [20]. And the checker is not fully stable against itself. Verdicts held on seven of eight pairs run twice, while the omission list changed on four of eight, which is why that list is offered as a prompt to look rather than a count [15].
None of this makes the report weak. It makes it readable by someone deciding what to buy, because both error classes are visible separately instead of averaged into a single number that hides which one would reach their patients.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Krasyn says its judge prompt allows faithful clinical translation to count as Supported, and that on this corpus it let clinical conclusions through under that heading; each is sensible, which is exactly why a signer would not catch them either.
Krasyn ships Note Check: paste a visit transcript and an AI scribe's note, and it labels each sentence Supported, Unsupported, Contradicted, Scaffolding or Unverified against the transcript, raises three pure-code flags, and lists facts the note left out. It never edits the note; the clinician reads the whole note and signs it.
Krasyn publishes no accuracy figure for Note Check because it has not been measured against a clinician-adjudicated reference set.
Krasyn's stated alternative is to run the tool on open data someone else has already labelled, publish the counts and the disagreements, and say which side looks right on reading.
Omi Health publishes medical-note-eval under the MIT license: 300 synthetic primary-care dialogues, SOAP notes written for each by named frontier models, and labels from a panel of LLM judges.
The 2025 Omi benchmark ships per-note counts of unsupported claims from three cross-family judges; the 2026 benchmark publishes per-writer totals only.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Specific and reproducible, but single-source and unadjudicated
The run is documented at an unusual level of detail — named engine and judge, temperature 0, preprocessing steps, an MIT-licensed public corpus, per-dialogue examples with quoted note text — and the findings are self-damaging, which raises credibility. It is nonetheless one first-party post with no independent replication, LLM judges on both sides with no clinician adjudication, overlap between the two label sets inferred rather than matched on statement text, six synthetic dialogues, only 18 of 36 pairs scoreable, and one of sixteen flags left uncharacterised.
Shipped and free, no external uptake evidence
Adoption evidence is limited to availability: Note Check exists as a production service, is free inside a Krasyn Scribe account, and was exercised in the live product for eight pairs. The supplied source discloses no user counts, no customer or clinic deployments, no third-party integrations and no usage of the tool by anyone outside Krasyn; the notes tested came from general-purpose models under a benchmark prompt rather than any commercial scribe. The referenced Omi corpus is publicly available but its own uptake is not quantified.
Understated: refuses a headline number and publishes its own misses
The framing runs opposite to typical vendor evaluation posts. Krasyn declines to publish any accuracy figure, reports that its checker matched the reference panel on only one of ten flagged notes, and details a deterministic rule that produced twelve wrong flags and a judge prompt that waved through unsupported clinical conclusions. Nothing in the source overstates capability relative to the evidence shown; the residual risk is the opposite — self-selected examples and a vendor-authored 'which side looks right' reading with no clinician adjudication, which is why the score is not more strongly negative.
Vendor-authored with product funnel, partly offset by against-interest disclosure
The sole source is written by the vendor about its own product, published on a developer platform, and closes with a UTM-tagged link ('utm_campaign=scribe-redteam-2026') to Note Check free inside a paid Krasyn Scribe account plus a pointer to Krasyn's own scribe faithfulness benchmark. The interpretive reading of which side was right is Krasyn's, and Krasyn also chose the six dialogues. Those distortion pressures are materially offset by the disclosure being self-damaging and by reliance on a third party's MIT-licensed labels, but no independent party checks any of it.
Moderate: detailed and self-critical, but wholly unverified externally
Confidence is supported by precision and direction — exact counts, named models, a stated engine and judge configuration, an open corpus, and findings that damage the author. It is capped by structural gaps: a single publisher and a single first-party source, no clinician adjudication, inferred overlap between label sets, six synthetic dialogues, half the pairs unscored, and no evidence of behaviour on real clinical encounters or of any external use.
build
The scribe drafts, the clinician verifies: an EMR vendor publishes its own faithfulness math1 distinct publisher
build
PerceptionBench puts a number on the step your pipeline treats as free1 distinct publisher
invest
Google says frontier models already know the facts they get wrong. That is a budget decision.1 distinct publisher
build
2,513 tool calls, zero refactorings: what agents actually do when you ask them to refactor1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 21, 2026