Build1 distinct publisher2 min readUpdated
One developer's evaluation layer argues that traces of successful sessions cannot say which documents helped, and that more traffic makes the wrong answer tighter rather than better.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
The interesting failure is not the missed document. A miss produces a bad outcome and gets fixed. The one that survives is the wrong material sitting beside a task that succeeded anyway, which then looks useful in the record afterward [1]. What the system does with that record is credit the document for work the person did [4].
Where the coin flip sits in the code decides whether it means anything. In ZCL the decision happens in the context provider, which creates the session id, decides whether the request should explore, asks value attribution for learned document values, handles cold start, and filters selected document ids through causal evidence before the bundle goes back to the agent; assignment happens ahead of that handoff [13]. A flip applied after selection would only relabel a choice the user's judgement had already made.
The price shows up in the sample budget. The default minimum sample threshold is 20, and it rides on the hypothesis object itself so the brake travels with the proposed change [10]. The generator builds order hypotheses from the top five documents already used for a task type and inclusion tests for up to three uncertain documents [11]. Three inclusion tests at twenty observations each is sixty sessions of a single task type before those questions can resolve, and one hundred and twenty if twenty is the count per branch rather than the total [15]. A deployment with real traffic absorbs that. A deployment serving one developer's agents spends months on three questions.
The author is direct about the ceiling. ZCL is described as his own context learning platform, and the post is offered as an account of how much the layer can currently prove, which he says is less than its own vocabulary suggests [8]. Whether the minimum sample threshold is connected to anything that blocks promotion is deferred to a later section [14]. That deferral is where the load sits. A metric declared in advance disciplines the analysis honestly. A threshold that no code enforces is a comment with a number in it.
None of this is a published result, and one repository is not evidence that randomized provisioning produces better agents. The diagnosis travels further than the implementation: when a memory system's only teacher is the record of tasks that went well, its teacher is the users who were going to do well [4].
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
An agent can fail because it missed the right material, or because it was given the wrong material with confidence; the quieter case is the wrong material sitting beside a successful task and then looking useful afterward.
A skilled user may ask for architecture notes because they know where to look, and those notes can travel with a good result even when they did not cause it; a system learning from that trace alone may send extra text that feels justified and still wastes the agent's attention.
The same expertise that makes someone request the right document also makes them likelier to finish the task without it, so skill is a common cause sitting upstream of the relationship being measured.
A system that reads the trace and credits the document has attributed to the context what belonged to the person.
There is no way to subtract the bias afterward from the trace alone, because the trace does not record why the material was requested.
The only cheap instrument that removes the bias is deciding who gets the material by coin flip instead of by request; randomization makes assignment independent of skill, so whatever difference survives between branches is attributable to the material.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Detailed self-report, no results or outside verification
The methodological argument is coherent and internally consistent, and the implementation description is specific enough to be checked in principle (file paths, object fields, min_samples=20, top-five/up-to-three generator limits, bundle bookkeeping). But everything comes from one self-published post by the system's own author, no experiment output, effect size or session count appears, the enforcement of the sample-size brake is explicitly deferred, and the source itself concedes the randomization design violates session independence. That caps evidentiary weight well below the midpoint.
Author's own agents only
The only adoption evidence is the author's disclosure that he runs ZCL for his own AI agents with the experimentation layer active during real work. No external users, installs, deployments, releases, downloads or third-party evaluations are reported anywhere in the supplied material, so adoption is near the floor rather than absent.
Experiment vocabulary slightly outruns what is shown
The framing — a missing evaluation layer, randomized assignment converting observation into experiment, differences 'attributable to the material' — is stronger than the demonstrated state, which includes no completed experiment, an unresolved sample-size brake, and admittedly non-independent session-level randomization. The overstatement is mild rather than severe because the author pre-empts it himself, stating the layer proves less than its own vocabulary suggests and naming the place the diagram lies.
Author writes about his own platform, discloses it
The post is written by the builder of ZCL about ZCL on a developer-marketing-friendly platform, which is a clear promotional incentive. It is partly offset by open disclosure of the relationship, a candid limits section, and admissions against interest (coverage deliberately dropped, brake possibly unwired, diagram misleading on independence). No pricing, funding, sponsorship or vendor relationship is disclosed either way.
Confident on the argument, thin on the artifact
Confidence is moderate: the reasoning about confounding and about pre-registering metrics is verifiable on its own terms and clearly stated, so claims about what the author argues and built are reliable. Confidence about whether the layer works as described is low — one self-interested source, no results, a truncated body, deferred enforcement, and an unresolved per-branch versus total sample ambiguity.
build
Your inference bill is an architecture defect: declare the task before you call the model1 distinct publisher
build
Netflix's plain-text recommender won on 40x fewer labels, and the bill moved rather than vanished1 distinct publisher
build
The reason your agent gets worse after an hour is that nothing ever leaves the context window1 distinct publisher
build
Agent reliability is a harness problem, not a prompt problem1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 23, 2026