Build1 distinct publisher2 min readUpdated
A replay that held the agent's actions fixed produced different labels before and after delayed operations resolved, and one late write moved the following run's score.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
The replay holds the agent's actions fixed and varies only the accounting [4]. That is what makes it a mechanism demonstration rather than a benchmark result: if the actions did not change, the divergence between the two labels cannot be charged to the model behaving differently [15]. Same trajectory, two labels, and the one that gets published depends on when the evaluator looked.
The second half of that finding is the part an infrastructure team can act on. The carryover into the next run appeared when service state persisted and did not appear after isolation or verified reset [c5b], which places the defect in the harness configuration rather than in the system under test [16].
The two conditions also do not come as a pair [3]. Polling until the write lands settles whether the write succeeded, and leaves the written state sitting where the next run can read it. A fresh container prevents that read, and tells you nothing about whether the operation the score refers to ever finished [3]. Each of the two obvious fixes buys exactly one condition.
Which is roughly why the protocol review found what it found. The endpoint is defined by the harness: the evaluator stops asking for actions after the model reports completion, or on a success condition, a timeout, or an action limit [7][8]. All ten protocols document that boundary and what gets scored [c6a], because all of it is in code the evaluator owns. The status of a pending write, a leftover credential, or an account created mid-run sits in a service the evaluator does not own [7]. So counting a stopped run as one trial [1] is an assertion about the environment, offered in place of evidence about it [13].
The two field reports the authors cite map cleanly onto the two conditions. The UK AI Security Institute's concurrent samples reusing accounts and artifacts left by other samples is separation failing [9]. The OpenAI and Hugging Face accounts of evaluation activity that used external services and continued past the original runtime is persistence past the endpoint, the case where the scored outcome was still moving [10]. The proposed open-effects record does not repair either one. It lists what may still be live after the endpoint, its status, and whether it could change the score or reach another run [14]. That is a disclosure format, and its whole value is that it gives a harness somewhere to write down what it does not know.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
In a controlled replay where the agent's actions were held fixed, the endpoint label and the terminal label differ for every delayed operation.
In a review of ten public protocols, all protocols identify when a run stops and what is scored.
In the same review of ten public protocols, unfinished operations and the evidence for treating runs as separate trials are documented less consistently.
Current agent evaluations score models on the state visible at the end of a stopped run, which they count as one trial.
Interpreting the end-of-run score as a final result would require two conditions that the endpoint does not itself necessarily establish: outcome finality and cross-unit separation.
The two conditions are independent: reconciling a delayed outcome can settle the label while runs still share state, and isolating runs can prevent carryover while the scored outcome remains unfinished.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Coherent mechanism demonstration, thin corroboration
The mechanism is demonstrated cleanly in principle: a replay with the agent's actions held fixed isolates read timing as the cause of label divergence, and the cross-run effect appears and disappears with state persistence, which is a well-formed controlled comparison. Against that, the entire cluster is one non-peer-reviewed arXiv preprint; the supplied text gives no replay scale, no named list of the ten reviewed protocols, and no coding methodology, and the two supporting field reports are secondhand citations whose primary accounts are not in the cluster. The authors themselves state the incident evidence does not establish frequency elsewhere.
No disclosed adoption of the proposed practice
The supplied source shows nothing about anyone adopting the completion argument or the open-effects record: no protocol has been revised, no harness has shipped the reporting artifact, and no maintainer has responded. The field incidents in the cluster evidence that the underlying problem occurs, not that the paper's remedy is in use, so adoption cannot be scored without inferring facts the source does not provide.
Scoped tightly, slightly under-claimed
The framing is unusually disciplined for a methodology preprint: the replay is labelled a mechanism demonstration, the incident reports are said to show boundary crossings without establishing frequency, and the recommended remedy allows retaining uncertainty rather than forcing a label. The one strong-sounding phrase — labels differed 'for every delayed operation' — is explicitly bounded to the controlled replay. If anything the practical consequence (that connected runs make reported run counts an upper bound on independent observations, so published agent scores may be less well-determined than they look) is stated more quietly than the evidence would allow, hence a mildly negative gap.
Standard academic framework-promotion incentive
The authors advance their own conceptual apparatus (the completion argument) and their own artifact (the open-effects record), and the ten-protocol review is framed as a gap that their proposal fills — a routine incentive to find documentation deficiencies. Offsetting this: the venue is a preprint rather than a product launch, nothing is sold or licensed, the closest prior work is credited at length in Section 1.1, and the limits of the evidence are stated by the authors themselves. The supplied text discloses no funding, vendor affiliation, or competing benchmark interest, so the assessment covers only the visible framework-promotion incentive.
Internally consistent, single unreviewed source
Confidence is moderate. The conceptual claims are clearly stated and mutually consistent, and the replay design supports its own causal reading, so the direction of the finding is credible. But there is one publisher, no peer review, abstract-level reporting of both empirical studies, unnamed protocols, secondhand incident attributions, a truncated body in the supplied text, and no adoption signal at all — enough missing detail that magnitude and generality cannot be judged.
build
OpenAI's president says open weights will accelerate the threat. His own cyber model stays gated.1 distinct publisher
build
19 unsanctioned actions in 10 of 122 runs: nothing escaped, and that is the point1 distinct publisher
leadership
Z.ai held back its own GLM-5.3 weights, and open-weight roadmaps have a new failure mode3 distinct publishers
build
Two rejected papers: the shadow evaluation that undercuts autonomous AI research claims1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 24, 2026