Build1 distinct publisher3 min readUpdated
Krasyn printed the definitions next to its hallucination rate and shipped a checker that runs on any scribe's note. The number it will not publish is anyone else's.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
The two figures the post leans on do not share an axis. Published evaluations put ambient-scribe hallucination near 1 to 3 percent of notes [3], while a March 2026 analysis of 71,173 AI-drafted and finalized note sections found a confirmed edit in 5.8 percent of them [4]. A note contains many sections, so the rates cannot be stacked [24]. Taken on its own, the second works out to roughly 4,100 sections that a human changed after the model was finished [5]. An edit is not proof of an invention, but it does mean somebody read closely enough to find something worth changing, and that reading is the part no product shipped.
Krasyn's answer is to define the unit before quoting a rate. An assertion is one atomic statement that could be true or false by itself, so "Denies fever, chills, and nausea" counts as three [6], and each one takes exactly one label of supported, inferred, unsupported or contradicted, with inference tracked apart as the contested case [7]. Hallucination rate is unsupported plus contradicted over all assertions, and coverage is scored separately against key facts, because a note reading only "Patient was seen" is perfectly faithful and worthless [8].
That second scale is where the report costs itself something. The production grounding verifier removed 12 sentences and strict coverage fell 1.4 points, all of the loss in one case where the judge threw out a sentence carrying a fabricated denial and a true hedged finding together [14]. Against 144 key facts [9], 1.4 points is about two facts [18]. The price of removing that fabrication was two things the clinician needed, and the defect is the granularity of the scalpel rather than the judgement behind it.
The trap layer is the piece worth copying whatever the score says. Traps are a regex plus a written rationale, so they cannot drift when a model changes, and one fires only if the mention survives negation and irrealis suppression scoped to its own sentence, with every suppressed mention logged beside the rule that killed it [10]. Five of the 61 fire in the grounded arm, two in the crosstalk case where the note correctly attributed the headaches to the spouse and the regex went off regardless [15]. Set those two aside and the residual is 4.9 percent [20]. They are left in deliberately, because a rule tight enough to silence them would also hide a real wrong-patient attribution [15]. A vendor tidying its own board deletes that trap instead.
There is a check on the checker: each run injects three unambiguous fabrications into every grounded note, which is 36 across 12 notes [19], and 36 of 36 were caught, with a sub-100 result defined in advance as grounds for declaring every other number suspect [13]. The weak joint sits elsewhere. gpt-4o produced the labels and also wrote the notes, self-preference bias in model judges is documented and uncontrolled here, no clinician has adjudicated a single label, and the harness author, an AI agent, wrote the corpus, the traps and the judge prompts [16]. By the post's own standard, that a benchmark without definitions is marketing [22], what survives is the definitions and the tool, not the figure. Krasyn publishes no competitor number, citing no API access and no matched corpus [11]. So the useful artifact ends up in the buyer's hands rather than the seller's: a transcript, a signed note, and a list of the sentences the transcript will not carry.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
The author founded Krasyn, an outpatient EMR with an AI scribe inside it, which has run a working outpatient clinic's real patient records since March 2026, so what the scribe drafts ends up in charts that real clinicians sign.
In August, Krasyn shipped two things: a published benchmark of how faithful its scribe's drafts are to the transcript, and Note Check, a tool that reads any scribe's note against its transcript and lists what the transcript does not support.
Krasyn measures at the level of a clinical assertion, one atomic statement about the patient that could be true or false on its own; "Denies fever, chills, and nausea" is three assertions, a measurement and its value are one, and hedging is kept verbatim.
Every assertion gets exactly one label against the transcript: supported (stated, or a faithful paraphrase or clinical translation), inferred (not stated but a reasonable clinical inference with a basis in the transcript, tracked separately because it is the contested category), unsupported (no basis at all), or contradicted (the transcript says the opposite, including a denied symptom, a declined treatment, or another person's symptom attributed to the patient).
Hallucination rate is unsupported plus contradicted over all assertions; coverage is measured separately against key facts per case because a note saying only "Patient was seen" scores perfect faithfulness, and a faithfulness gain bought by dropping content is treated as a regression.
The corpus is twelve synthetic transcripts, 144 key facts and 61 traps, all original invented dialogue with no real or de-identified patient data; seven of the twelve are adversarial, and the corpus was committed to git before the harness existed because the repo has a documented habit of expectations written to match current output.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Transparent method, thin and self-judged data
The methodological disclosure is above the norm for vendor benchmarks: printed definitions, a model-free regex trap layer, corpus committed before the harness, paired arms from one generation, and a per-run negative control with a stated invalidation rule. But the empirical base is twelve synthetic written-English transcripts, the labels come from the same model that wrote the notes with self-preference bias explicitly uncontrolled, no clinician has adjudicated any label, and the whole thing is one self-published post with no external replication. The external baselines it leans on are cited without attribution and use mismatched denominators.
One clinic, no third-party uptake shown
Evidenced adoption is one working outpatient clinic in production since March 2026, plus two August releases. Note Check names nine scribe products it can read as pasted text, but that is stated capability, not usage: no user counts, download figures, integrations, or external adopters of the benchmark or the checker are disclosed anywhere in the material.
Claims run behind the disclosed work
The post consistently claims less than its own material would allow: it declines to publish any competitor number for lack of API access and a matched corpus, foregrounds the coverage cost of grounding, reports traps that still fire including false positives it deliberately leaves in, and volunteers that judge self-preference is uncontrolled, that no clinician has adjudicated a label, and that fabricated denials evade both layers. The offsetting overstatement is minor: an unattributed market framing and a single-clinic deployment described in language that reads broader than one site.
Vendor grades own product, markets checker to rivals' users
The author is the founder of the product being measured, publishing on a developer platform under the company handle. The benchmark scores only Krasyn's scribe, and the companion tool is pitched at clinicians using nine named competitors while carrying no accuracy measurement of its own. That is a direct commercial interest in both the number and the distribution channel. The countervailing signal is that the same post withholds competitor figures and publishes its own defects, which cuts against pure promotion.
Rich internal detail, zero external corroboration
Confidence is moderate. What Krasyn says it does is described in enough operational detail to be checkable, and the self-reported limitations are specific rather than boilerplate, which raises trust in the account of the method. But every figure comes from one self-published vendor post with no independent replication, no clinician review, a twelve-case synthetic corpus, and unattributed external baselines, so confidence in the magnitude of the faithfulness results stays well below confidence in the methodology description.
build
Pick a log anomaly detector on volume, latency and secrets, not on which one is smarter1 distinct publisher
build
Four months of A100 bills say self-hosting is a utilization bet, not a cost saving1 distinct publisher
leadership
Re-baseline AI procurement on cost per completed task, not dollars per million tokens1 distinct publisher
build
Deferred tool schemas cut cost 21% on average, and made one task type 12.3% dearer1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 21, 2026