Build1 distinct publisher3 min readPublished
The pooled 17% belongs to a table with a 40% outlier in it, so the part that transfers is the scoring rule: a mechanical null check against fields someone has verified are missing, plus a blank-rate column to keep it honest.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Pull the outlier and the headline moves. If each of the six setups faced the same 96 absent fields [5], the pooled 17% [8] is a straight average, and removing the 40% row [7] leaves (17 x 6 - 40) / 5, about 12.4% [1]. Velrim's own 11% [6] sits inside that band. So 17% is a property of a table with Mistral in it. A stack that does not include Mistral should carry a prior nearer 12%.
The finding also rests on a small ledger. Six systems against 96 absent fields is 576 chances, and 17% of that is roughly 98 invented values [3]. Per system, 11% is about eleven fields and 40% is about thirty-eight [2]. Velrim says as much in its own framing: where the uncertainty brackets overlap, the test cannot separate two systems [24], and per document type, gaps under about 4 to 10 points are invisible at this corpus size [17]. Read the table as one outlier and five rows that tie.
One thing does not reconcile. The answer keys are described as holding 96 absent fields [5], while the team also says it hand-checked all 142 absent-field labels against the actual pages before any competitor system ran [23]. That is a 46-label gap [5] the post does not explain, so count the keys in the repo before you quote a rate. The sequencing is right either way: labels frozen first, competitors second.
For the rate to mean anything in your pipeline, your documents have to look like those 124 [4] and your schema has to ask for absent fields at a similar frequency. The measurement is per absent field, so your exposure across all returned fields is the absent share multiplied by the fabrication rate. A schema of mostly required fields carries very little of it. An optional-heavy schema carries nearly all of it. Velrim's weakest accuracy row is TV ad contracts at 0.59 out of 1 [16], which tells you which document families are hard, not what your own score will be.
The decoding ablation is the part I wanted more of. Constrained and structured runs force the output to match the schema as the model writes it, while free-decode runs write freely and get checked afterwards, and on this corpus that choice is worth under 2 points for OpenAI's model and nothing detectable for Gemini [18]. That comparison is reported on accuracy.
The pre-registration is the craft worth copying. Section 10 of the committed plan states that the expected accuracy outcome between Velrim and the bare model underneath it is a statistical tie, signs that tie as publishable before any money is spent, and adds that a design needing Velrim to win the F1 column is a failed design [12]. The run came in two points behind bare Gemini, a statistical tie [13]. The post also reports a DIY setup on OpenAI's small model beating Velrim by about four points [14], and bare Gemini beating it clearly on US registration filings [15]. A vendor paying for a benchmark that ends with "if a wrong extracted field costs you nothing, the bare model is a much better purchase" [25] is not the usual genre.
What is reusable here is the answer key with verified absences in it, plus the balancing column that stops refusal from being a winning strategy [22]. Without both, an extraction suite scores the fields that exist and stays quiet about the ones that do not [3].
Ranked by verification strength, evidence, and original report placement.
Velrim ran the comparison and sells one of the six APIs in the published table.
Every raw output, every request ID and the scoring CLI used for the comparison are published in the velrimhq/velrim-eval repository.
124 real documents were sent through six setups, which were asked to extract fields that are not present in those documents.
The answer keys hold 96 absent fields, for which the correct answer is null, meaning the field does not exist.
Velrim's own system invented answers for about 11% of the absent fields, and the other systems were about the same, with differences within noise.
The Mistral setup invented values for about 40% of the missing fields, the only system outside the cluster.
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · September 2, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
build
A 14,000-star watermark remover, and no detector to test it against1 distinct publisher
product
Four leaderboards, four denominators: what you buy when you standardize on a coding agent1 distinct publisher
build
A RAG demo becomes a product at the tenant boundary, not the retriever1 distinct publisher
build
The Tokenizer Is Your Real Price List, Not the Per-Million Rate Card1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Self-published, unusually checkable
Every number here — the 11%, the 40% outlier, the pooled 17%, the price multiple — comes from the company being measured. What keeps it from being a brochure is specific and verifiable: a plan committed before the first paid API call, request IDs and raw outputs in a public repository, a judge that is a string comparison rather than a model, and 142 third-party absent-field labels re-checked by hand against the pages. What holds the score down is equally plain: nobody outside Velrim has re-run it, the rates rest on about a hundred individual field decisions, and Velrim itself says per-document-type gaps under 4 to 10 points cannot be read.
Published, not yet used by anyone else
One thing has actually happened: on 2 September 2026 Velrim put the run and its artifacts online. No independent party has scored a system against the null-field rule, none of the five compared vendors has responded, and the two products Velrim says also ship per-field confidence could not be tested at all because their terms forbid benchmarking. This is a method with a repository, not a method with users.
Undersold, except in the headline
The weakest thing in this report is its own top line. A pooled 17% only reaches 17% because one setup invented values for about 40% of absent fields; the other five sit near 12%, and the summary line says so while the title does not. Everything else leans the other way. The founder writes that the 7-10x premium 'doesn't buy you accuracy', prints a four-point loss to a do-it-yourself OpenAI setup and a clear loss to bare Gemini on US filings, names 0.59 as his worst document type, and volunteers that his caution leaves three times as many present fields blank as anyone else. A vendor write-up that spends more ink on where it loses than where it wins is claiming less than its own data would allow.
Grading its own table, and saying so
The conflict is stated in the first sentence and then reinforced: Velrim sells one of the six APIs, and the founder says he commissioned the whole run because he could not justify charging 7 to 10 times the raw models. Velrim also picked the corpus, wrote the answer keys, and removed 46 of 142 absent-field labels from the count. The pre-signed prediction of a tie, the per-label record in the repository and the published CLI are genuine constraints on that discretion — they make bad faith detectable — but they do not transfer it to anyone else.
Internally consistent, externally untested
We are reading one publisher, one author, and that author is the vendor. Internal consistency mostly holds under pressure: the apparent contradiction between 96 absent fields and 142 hand-checked labels dissolves once you reach the 40 wrong and 6 unresolved labels that were struck for every system. But consistency is not verification, the sample is thin enough that Velrim itself flags an unreadable range, and until someone runs the published CLI our confidence has nothing outside Velrim's own account to lean on.