Build1 distinct publisher3 min readPublished
A dev.to argument for lifting numbers out of documents by pattern before the model ever sees them rests on a cost asymmetry between gaps and wrong digits, and on a schema guarantee that covers shape rather than truth.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
Follow any of these and your For You feed starts watching them — no settings page required.
build
An unsupervised agent loop billed $38 before anything in the system said stop1 distinct publisher
build
244 kB, 500 a minute, 5 percent: three ceilings that fail for the same reason1 distinct publisher
build
Your inference bill is an architecture defect: declare the task before you call the model1 distinct publisher
build
OpenClaw makes the channel the architecture, and the reasoning loop a lodger1 distinct publisher
Assignment is a narrower job than reading, and in the dev.to design the narrowing is the whole point. The model receives the stored text plus candidate values that a pattern pass already lifted out, and it returns a mapping rather than a value [20]. Errors then have to surface as a figure in the wrong slot instead of a slightly wrong figure in the right one, and only the first kind gets noticed [1].
Do the arithmetic on the author's own preference [7]. Take a thousand fields. The tolerant system leaves eighty of them empty, at roughly a minute of attention each, so about eighty minutes of clerical work [4][3]. The tidy system fills all thousand and is quietly wrong about twenty, and at the piece's own propagation figure of three copies apiece, that is eighty records carrying a wrong number [5][4]. That is the trade in plain terms: eighty minutes of scheduled work with a named owner against eighty records that need reconciling once somebody has already quoted one of them.
The edit-distance claim is checkable with a calculator. 349,000 to 340,000 is one character and nine thousand in value [2][1]. 132 square metres to 138 is also one character, and the author prices that gap at nine thousand euros, which implies about 1,500 euros per square metre [11][2]. Whether that implied rate matches your market does not matter; the ratio between how wrong the string looks and how wrong the money is stays lopsided.
The mechanism offered is digit-level. The cited research has models encoding numbers digit by digit in base 10, with errors that sit close in string edit distance and far away in value [9]. The author flags the gap rather than papering over it: those papers measure arithmetic and counting, not copying out of a document [10]. For the finding to transfer, the copy path would have to run through the same digit-level representation instead of a span lifted verbatim from the input, and nobody has shown that here. What stands without it is the weaker claim, that an error in a number need not look like an error, and the weaker claim is enough to justify the pattern pass [12].
The guarantee most teams already bought covers shape. OpenAI's documentation promises the response adheres to the supplied schema, with no missing required key and no invalid enum value [13], and the same documentation says structured outputs can still contain mistakes and that input unrelated to the schema can still produce hallucinations [14]. A price key will exist and will hold a number, but the schema guarantee stops there and does not check that number against the document [15]. A well-formed row looks audited whether or not it is.
Adoption cost lands on document variety. The candidate pass is very reliable on clean HTML and on PDFs with an embedded text layer [21]. On a scan the uncertainty starts at text recognition, and tables, differing decimal separators and currencies degrade the candidate list before the model is involved [22]. So the pattern-first layer does not fix scans; it removes one failure path and inherits another. In my context, HTML and text-layer PDFs feeding figures that get quoted downstream, that is the right trade, and the stored copy is the cheap half of it [17]. For photographed faxes, the honest answer is a confidence gate and a person on the digits.
Ranked by verification strength, evidence, and original report placement.
The dev.to piece asserts that a missing number gets noticed while a wrong one does not.
According to the piece, when a model turns 349,000 into 340,000 the result is a plausible number, in the right field, in the right format, and nobody has any reason to go look it up.
An empty field is described as an interruption: someone sees the gap, opens the source and fills it in, at a cost of a minute of attention, and the system stays trustworthy.
A wrong field is described as a decision that gets forwarded, quoted in an offer and added to a total; by the time anyone notices it has been copied into three other places, leaving four versions in circulation.
The author notes those papers study arithmetic and counting, not copying out of a document, so they do not prove that every extraction error arises this way.
The piece gives a second example: a property of 132 square metres becomes 138, which looks like a typo and behaves like nine thousand euros.
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 31, 2026
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One essay leaning on one vendor doc
The only claim here that an outsider can check is OpenAI's, quoted in both directions: the schema is honoured, and the contents can still be wrong. Everything about how numbers break inside a model arrives as unnamed research, and the sweeping line — most extraction projects fail quietly in the third digit — has no project, log or rate behind it. The arithmetic in the examples holds up, but it is arithmetic on the author's own hypotheticals.
Nothing shipped to point at
No release, no repository, no client system, no benchmark and no volume figure — only a described five-layer arrangement and the author's statement that he no longer trusts the alternative. There is no way to tell whether this pattern runs anywhere beyond his own work.
Modest voice, unmeasured cure
Slightly overstated, and less than you would expect from the genre. The author repeatedly clips his own claims — the papers do not prove the mechanism, deterministic means traceable rather than correct — which is rare and pulls the gap toward zero. What pushes it back up is the shape of the argument: a vivid, undocumented failure rate on one side and a five-layer remedy with no measured outcome on the other, with 'very reliable' doing the work a number should.
Methodology post under a studio byline
This is a practitioner publishing a build pattern under a studio handle on a platform where reputation is the currency, and nothing about a client, product or dataset is disclosed — the ordinary pull of methodology writing, which flatters the method its author already uses. Cutting the other way: the one external authority he leans on is a vendor's admission against its own feature, not a favourable citation, and no tool or service is offered for sale anywhere in the piece.
Reasoning checkable, results not
We can be fairly sure of what the argument says and where it is thin, because the logic is inspectable and the numbers in the examples reconcile. We cannot be sure the world behaves as described: one voice, no second account, and no observed outcome from running the pattern. Read it as a well-reasoned hypothesis about a real failure mode, not as a finding.