Build1 distinct publisher3 min readPublished
The comparison buyers of AI data-governance tooling get shown is governed metadata against full-context stuffing; this harness also ran it against a cheap keyword filter, four times, under contracts frozen before collection.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
The rule that keeps failing is a ratio, and that is the whole mechanism. At the R5 holdout the governed route sent 763.5 prompt tokens and the lexical route sent 110.0, which the author records as 6.94x against a frozen ceiling of 3.0x [7]. Passing would have meant fitting the same selection into 110.0 x 3.0 = 330 tokens, a cut of 57 percent on what governance actually sent [15]. Nothing in that arithmetic accuses the governed route of waste in absolute terms. 763 prompt tokens is not a large prompt.
The denominator moved for a defensible reason. R3's catalog was lexically tractable, so a keyword filter had real signal to match, scoring 0.660, then 0.737, then 0.631 mean F1 as the catalog grew [5]. Rebuilding R4 and R5 on opaque physical names removed that signal [6]. A route that matches nothing also retrieves nothing to describe, so its prompt shrank to 110 tokens, and under a ratio ceiling the denominator shrank with it [7]. My read: cost per correct answer is the metric this comparison wants, and it is undefined for a baseline scoring zero.
The other figure worth pricing is the direction of travel. Governed mean F1 went 0.780, then 0.632, then 0.447 across the three catalog sizes at the prespecified 0 percent classifier-miss condition [5], a relative fall of 42.7 percent from the smallest catalog to the largest [17]. That is the opposite of the scaling behaviour a catalog gets bought for.
Now what would have to be true for the sponsored number to transfer. The McKnight Consulting Group study this harness reproduces in structure, "Stop the Token Bleed" by Jake Dolezal and William McKnight, held the model constant and reported up to roughly 89x fewer tokens at scale with F1 1.000 against 0.29 to 0.66 ungoverned; it was sponsored by Informatica, a Salesforce company [11]. The baseline there is ungoverned stuffing, and this harness agrees on that comparison [10]. So the 89x is a claim about shops whose current practice is to shovel the schema into the window. Where a keyword prefilter already sits in the pipeline, that study did not run your experiment.
The earliest rejection was on quality rather than cost. At R2.1's 3,000-object holdout the governed route scored 0.24065 against a prespecified floor of 0.245533, short by about 2 percent, while the lexical prefilter scored 0.588, or 2.44 times the governed score [13][18][19]. The author discloses that he is employed by Salesforce and publishes the harness so readers can generate their own numbers instead of trusting his [12]. R3 and R5 ship public decision packs; R4's pack is held, and labelled as held [14].
Ranked by verification strength, evidence, and original report placement.
The author built a benchmark family to test whether a governed metadata layer earns its cost when an agent selects enterprise context, and ran it four times under four frozen contracts, redesigning the catalog and changing the acceptance ceiling along the way.
The preregistered claim that governance earns its cost against a cheap baseline was rejected in every round that tested it (R3, R4, R5); the earlier R2.1 round was rejected on a different rule.
R3 compared three routes on the same local model: raw full-context stuffing, a cheap lexical prefilter, and a governed metadata route.
At the R3 holdout the keyword filter beat the governed route on the observed means and used less than half the prompt tokens; the paired F1 interval was [-0.371, 0.00005], which does not exclude zero in governance's favour, and the token ratio was 2.11x against a frozen ceiling of 1.10x.
Mean F1 across 20 seeds at the prespecified 0% classifier-miss condition: the governed route degraded monotonically as the catalog grew, at 0.780, 0.632 and 0.447, while the lexical route scored 0.660, 0.737 and 0.631.
R3 used a lexically tractable synthetic catalog; R4 and R5 moved to semantic-access catalogs built on opaque physical names, and in R5 the lexical route scored 0.000 F1 at every catalog size.
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 28, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
invest
Salesforce's double digits, minus Informatica: agentic AI is real and still 2% of revenue1 distinct publisher
product
Salesforce's longer-dated backlog grows at half the rate of its cRPO headline1 distinct publisher
invest
Salesforce bought $27.1B of its own stock in a quarter it grew 13%2 distinct publishers
invest
Salesforce's 14% cRPO is the best agent-demand read yet. Organic growth is 6.5%2 distinct publishers
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Rigorous method, no second witness
The unusual thing here is the order of operations: acceptance rules written before the numbers, a 0% classifier-miss condition specified in advance, 20 seeds per cell, a paired F1 interval reported as [-0.371, 0.00005] rather than a bare mean, and decision packs published for R3 and R5. That is stronger discipline than most vendor benchmarks get. What holds the score down is corroboration, not care — one author, one local model, catalogs he designed himself, and R4's pack held back.
One person's runs, invitation open
Everything observed here was run by the same hands: R2.1, R3, R4, R5, twenty seeds a condition, on a local model. The author asks explicitly to be contradicted and nothing in this reporting shows anyone taking him up on it. No customer, no deployment, no named commercial layer, and no sign that the sponsored study's side has re-run anything.
Argues against itself, then generalises
The loudest sentence in this story is the author rejecting his own preregistered claim, alongside the admission that he relaxed the cost ceiling from 1.10x to 3.0x and then missed it at 6.94x. Writers protecting a position do not volunteer that. The one place the framing outruns the measurement is scope: a keyword prefilter undercutting governed metadata across four rounds is four rounds of one synthetic catalog family on one local model, and R5 shows exactly how completely a catalog redesign can decide which route has any signal at all.
Disclosed, and pointing the other way
Follow the money: the study being reproduced was paid for by Informatica, Informatica belongs to Salesforce, and the author states that Salesforce employs him. Every structural pressure in that chain favours confirming the 89x token saving — and the published result rejects his own cost claim in all three rounds that tested it, on catalogs he concedes were redesigned in a way that favoured his method on quality. The pressure is real, named up front, and contradicted by the output. What we cannot measure is the other side of the table: nothing tells us how Informatica or McKnight Consulting Group responded.
Recomputable, but single-channel
Two things cap this and neither is sloppiness. The ratios, intervals and seed counts can be recomputed as printed — 110.0 against 763.5 really is 6.94x — but they all arrive through the one person who ran them, as do the sponsored study's own figures. And R4, the round that ended inconclusive on a token-window overrun, is the round whose decision pack is disclosed as held rather than published, so it is the least inspectable part of the record.