Skip to content

Build1 publisherNot yet confirmed elsewhere3 min readPublished

The scribe drafts, the clinician verifies: an EMR vendor publishes its own faithfulness math

Krasyn printed the definitions next to its hallucination rate and shipped a checker that runs on any scribe's note. The number it will not publish is anyone else's.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • Krasyn published a benchmark of its AI scribe's faithfulness to the transcript in August, alongside Note Check, which grades any scribe's note against its own transcript.
  • The EMR has carried a working outpatient clinic's real patient records since March 2026, so the scribe's drafts reach charts that clinicians sign.
  • The test set is twelve invented transcripts, 144 key facts and 61 traps, seven cases adversarial, committed to git before the harness that scores them existed.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • capability A clinic can now point a fabrication check at the scribe already installed in the building, using its own transcripts, without asking that vendor for access or cooperation.
  • constraint Because there is no competitor arm, the published rate binds only Krasyn; anyone who wants a comparison has to build and defend the matched corpus themselves.
  • cost The safety step is paid for in deleted true content, and the bill lands on the clinician reading a note missing a real hedged finding that shared a sentence with a fabrication.
  • contradiction The post argues undefined benchmarks are marketing, then hands labelling to the model that wrote the notes, which makes the definitions and the checker the durable part and the score the softest.

The two figures the post leans on do not share an axis. Published evaluations put ambient-scribe hallucination near 1 to 3 percent of notes [21], while a March 2026 analysis of 71,173 AI-drafted and finalized note sections found a confirmed edit in 5.8 percent of them [22]. A note contains many sections, so the rates cannot be stacked [17]. Taken on its own, the second works out to roughly 4,100 sections that a human changed after the model was finished [23]. An edit is not proof of an invention, but it does mean somebody read closely enough to find something worth changing, and that reading is the part no product shipped.

Krasyn's answer is to define the unit before quoting a rate. An assertion is one atomic statement that could be true or false by itself, so "Denies fever, chills, and nausea" counts as three [3], and each one takes exactly one label of supported, inferred, unsupported or contradicted, with inference tracked apart as the contested case [4]. Hallucination rate is unsupported plus contradicted over all assertions, and coverage is scored separately against key facts, because a note reading only "Patient was seen" is perfectly faithful and worthless [5].

That second scale is where the report costs itself something. The production grounding verifier removed 12 sentences and strict coverage fell 1.4 points, all of the loss in one case where the judge threw out a sentence carrying a fabricated denial and a true hedged finding together [11]. Against 144 key facts [6], 1.4 points is about two facts [19]. The price of removing that fabrication was two things the clinician needed, and the defect is the granularity of the scalpel rather than the judgement behind it.

The trap layer is the piece worth copying whatever the score says. Traps are a regex plus a written rationale, so they cannot drift when a model changes, and one fires only if the mention survives negation and irrealis suppression scoped to its own sentence, with every suppressed mention logged beside the rule that killed it [7]. Five of the 61 fire in the grounded arm, two in the crosstalk case where the note correctly attributed the headaches to the spouse and the regex went off regardless [12]. Set those two aside and the residual is 4.9 percent [20]. They are left in deliberately, because a rule tight enough to silence them would also hide a real wrong-patient attribution [12]. A vendor tidying its own board deletes that trap instead.

There is a check on the checker: each run injects three unambiguous fabrications into every grounded note, which is 36 across 12 notes [18], and 36 of 36 were caught, with a sub-100 result defined in advance as grounds for declaring every other number suspect [10]. The weak joint sits elsewhere. gpt-4o produced the labels and also wrote the notes, self-preference bias in model judges is documented and uncontrolled here, no clinician has adjudicated a single label, and the harness author, an AI agent, wrote the corpus, the traps and the judge prompts [13]. By the post's own standard, that a benchmark without definitions is marketing [16], what survives is the definitions and the tool, not the figure. Krasyn publishes no competitor number, citing no API access and no matched corpus [8]. So the useful artifact ends up in the buyer's hands rather than the seller's: a transcript, a signed note, and a list of the sentences the transcript will not carry.

What to watch

  • Whether the judge agreement figure on the fixed 163-assertion list holds steady between runs, and whether it is reported when it moves.
  • Whether a clinician ever adjudicates the labels, which is the only step that settles the self-preference question.
  • Whether any buyer runs Note Check against a commercial scribe and publishes the result, since Krasyn will not.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories