Build1 distinct publisher2 min readUpdated
A deterministic fixture shows structure, provenance and dual-encoder agreement all passing on a handoff that described the wrong Git tree. One check read the repository and refused it.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
The row count is the part worth doing arithmetic on. The handoff recorded 12 rows where the tree held 1843 [6], wrong by a factor of about 154 [7], and none of the seven passing checks was built in a way that could notice it, because false content does not become malformed merely by being false [9].
Each of those checks makes a narrower statement than the word "verified" implies. Provenance confirmed that four objects reached the declared root at repo@a1b2c3d4 through connected edges [4]; what it cannot confirm is that the producer ever looked at that root, and a fixture that is consistently wrong about its branch produces a clean graph [10]. The checkpoint confirmed that the recorded state had not silently changed since it was committed [4], which is integrity, and integrity protects a false value from mutation exactly as well as a true one [11]. Authority agreement had three encodings agree [4]. The author's reading of that result is "this artifact has one stable interpretation," not "this interpretation describes what happened" [13], because both encoders read the same handoff and unambiguous wrong values survive re-encoding intact [12].
The fixture is filed as examples/common-mode-handoff.json [3], naming the mechanism rather than the symptom. Counting the printed output, the run holds eight checks: seven that read the artifact and one that reads something the producing agent could not have written [17]. That ratio is the finding. Six of the seven could be replaced with better versions of themselves and the wrong branch would still pass, because the thing they disagree about is the document, and the document is internally fine.
The repository check does not end the regress. According to the author it can inspect the wrong checkout, lean on a stale test result, or encode its own mistake, and what it actually does is move the trust root to a smaller, named checker rather than remove it [15]. That is a better place for the root, and it is still a place, which means the useful way to describe a verification suite is not how many checks pass but how many of them consult evidence the producer had no hand in. In this run the answer is one [14].
Worth keeping in view: the author states plainly that no language model produced this result and that the run says nothing about how often agents read the wrong branch [8]. The claim on offer is about what a passing suite is capable of asserting, and it holds whether the mistake rate is one in ten or one in a million.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
In a deterministic fixture, a producing agent reads the wrong Git branch and records an authentication provider and a row count that are both wrong for the working tree being handed off.
The handoff JSON is valid, its provenance graph is connected, its commitments recompute, and two separately implemented encoders produce the same semantic world.
The verifier output for examples/common-mode-handoff.json reads FAIL, with the summary line '7 verified, 1 FAILED'.
The seven verified checks were: structure (contract 0.1, 4 objects), identity (agent-b to agent-c), checkpoint (cp-4412-02), provenance (4 objects to repo@a1b2c3d4), retained constraints (2 MUST, 1 SHOULD), conflicts (none, 1 open), and authority agreement (3 encodings agree).
The only failing check was external truth, reported as EXTERNAL_RECEIPT_REJECTED: the repository working tree at repo@a1b2c3d4 rejected the world described by the handoff.
The rejection listed two discrepancies: auth.provider is okta-oidc in the tree, not auth0-oidc, and legacy.sessions counted 1843 rows, not 12.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Reproducible fixture, single self-authored source
The core demonstration is concrete and checkable: a pasted verifier transcript with named checks and values, a public 0.2.0 package with an offline demo, and explicit scoping that no language model was involved. It is nonetheless one item from one publisher, written by the tool's author, with no independent replication of the transcript and no external data on how often the failure occurs in real systems.
Shipped package, no usage evidence
The only adoption fact in the cluster is that the verifier exists publicly at version 0.2.0 with an offline demo and public fixture. No downloads, deployments, third-party users, integrations or production usage are reported anywhere in the supplied source, so adoption is barely above zero on release existence alone.
Claims stated narrower than the demonstration
The post consistently understates rather than overstates: it disclaims production relevance and frequency, notes no language model was involved, refuses the larger sentence agreeing encoders might seem to earn, and concedes the external receipt can itself inspect the wrong checkout or encode a stale result. The scoped tool description ('not an arbitrary-text truth checker or a hallucination detector') and the author disclosure further limit the reach. Slightly negative rather than strongly so, because the underlying evidence is a hand-built fixture and cannot support broader conclusions anyway.
Author-built tool, disclosed
The post promotes software the author wrote, closing with an install command and demo invocation, and the AI editing assistance is also disclosed. That is a clear promotional incentive, but it is stated openly, the package and fixture are public and runnable, and the tool's scope is deliberately narrowed rather than inflated, which partially offsets the incentive.
High on the artifact, low on generalisation
Confidence is high that the described transcript and layer-by-layer limitations hold as stated, because the fixture is deterministic, the output is quoted and the package is public. Confidence is low that this generalises to real agent systems or to any incidence rate, since the cluster has one self-interested source, no independent verification and no field usage data.
build
Three manual interventions in a month, and every guard was working as designed1 distinct publisher
build
Six MariaDB versions, one real difference: the only reason to leave 10.6 is the July 2026 clock1 distinct publisher
build
Force the tool call, then hand Lightsail a long-lived key1 distinct publisher
build
AI-written code fails the same four ways, and every gate you own reports green1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 22, 2026