BuildNot yet confirmed elsewhere1 publisher2 min readPublished
Ten agent eval protocols say when a run stops. Fewer say whether the result is settled.
A replay that held the agent's actions fixed produced different labels before and after delayed operations resolved, and one late write moved the following run's score.
The Engineer · Build desk
What happened
- An arxiv paper argues an end-of-run score counts as a final result only if outcome finality and cross-unit separation hold, and stopping a run establishes neither.
- In a controlled replay, the label read at the endpoint disagreed with the label read after the operation resolved in every delayed case.
- A write applied after the endpoint changed the next run's score whenever service state carried over between runs.
- Across ten public protocols, the inconsistently documented parts were unfinished operations and the evidence that runs are separate trials.
Why it matters
- constraint Once a route between runs is admitted, the run count becomes a ceiling on independent observations rather than the observation count, so error bars and win margins computed from trial counts are...
- exposure Anyone reusing a published number from a shared-service harness inherits a contamination the number has no field to declare, including procurement teams and safety reviewers who never touch the...
- decision Evaluators now face a three-way call on runs with operations still in flight: re-score after resolution, bound the effect, or publish the trial as uncertain instead of as a pass or a fail.
- contradiction The paper leans on incident reports to show the boundary really is crossed in production, then rules out using them to size the problem, which supports new record-keeping but not discounting any...
The replay holds the agent's actions fixed and varies only the accounting [1]. That is what makes it a mechanism demonstration rather than a benchmark result: if the actions did not change, the divergence between the two labels cannot be charged to the model behaving differently [17]. Same trajectory, two labels, and the one that gets published depends on when the evaluator looked.
The second half of that finding is the part an infrastructure team can act on. The carryover into the next run appeared when service state persisted and did not appear after isolation or verified reset [c5b], which places the defect in the harness configuration rather than in the system under test [18].
The two conditions also do not come as a pair [6]. Polling until the write lands settles whether the write succeeded, and leaves the written state sitting where the next run can read it. A fresh container prevents that read, and tells you nothing about whether the operation the score refers to ever finished [6]. Each of the two obvious fixes buys exactly one condition.
Which is roughly why the protocol review found what it found. The endpoint is defined by the harness: the evaluator stops asking for actions after the model reports completion, or on a success condition, a timeout, or an action limit [9][10]. All ten protocols document that boundary and what gets scored [c6a], because all of it is in code the evaluator owns. The status of a pending write, a leftover credential, or an account created mid-run sits in a service the evaluator does not own [9]. So counting a stopped run as one trial [4] is an assertion about the environment, offered in place of evidence about it [15].
The two field reports the authors cite map cleanly onto the two conditions. The UK AI Security Institute's concurrent samples reusing accounts and artifacts left by other samples is separation failing [11]. The OpenAI and Hugging Face accounts of evaluation activity that used external services and continued past the original runtime is persistence past the endpoint, the case where the scored outcome was still moving [12]. The proposed open-effects record does not repair either one. It lists what may still be live after the endpoint, its status, and whether it could change the score or reach another run [16]. That is a disclosure format, and its whole value is that it gives a harness somewhere to write down what it does not know.
What to watch
- Whether any of the ten reviewed protocols adopts an open-effects record, or starts publishing pending-operation status alongside scores.
- Whether anyone measures how often endpoint and terminal labels diverge outside a controlled replay, which would put a rate on the mechanism.
- Whether leaderboards begin asking for evidence of isolation or verified reset before accepting a reported trial count as a sample size.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence46
- Adoption
- Insufficient
- Hype gap−8
- Incentives34
- Confidence50
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
In a controlled replay where the agent's actions were held fixed, the endpoint label and the terminal label differ for every delayed operation.
- [2]
In a review of ten public protocols, all protocols identify when a run stops and what is scored.
- [3]
In the same review of ten public protocols, unfinished operations and the evidence for treating runs as separate trials are documented less consistently.
- [4]
Current agent evaluations score models on the state visible at the end of a stopped run, which they count as one trial.
- [5]
Interpreting the end-of-run score as a final result would require two conditions that the endpoint does not itself necessarily establish: outcome finality and cross-unit separation.
- [6]
The two conditions are independent: reconciling a delayed outcome can settle the label while runs still share state, and isolating runs can prevent carryover while the scored outcome remains unfinished.
- [7]
In the replay, a delayed write changes the next run's score when service state persists between runs.
- [8]
The delayed write did not change the next run's score after isolation or verified reset.
- [9]
The endpoint is the moment the evaluator stops requesting actions; a tool call may still be running at that point, and a file, credential, account or service value may remain available to later runs.
- [10]
During a run the model acts through the available tools until it reports completion or reaches a success condition, timeout, or action limit; that condition is the stop rule.
- [11]
The UK AI Security Institute reported concurrent evaluation samples finding and reusing accounts and artifacts left by other samples.
- [12]
OpenAI and Hugging Face separately described evaluation activity that used external services and continued beyond the original runtime.
- [13]
The authors state that these reports show effects crossed the nominal evaluation boundary in those settings, but do not show how often the same problem occurs elsewhere.
- [14]
The paper argues a final label is justified only when anything that could still change the claimed outcome is resolved, bounded, or retained as uncertainty.
- [15]
Treating runs as separate trials requires that no relevant route connects them; when a connection remains, the analysis must represent it or group the connected runs.
- [16]
The paper proposes an open-effects record listing operations or resources that may remain relevant after the endpoint, their current status, and whether they could change the scored outcome or affect another run.
- [17]
Because the agent's actions were held fixed in the replay, the difference between endpoint and terminal labels is attributable to when the outcome was read rather than to any change in model behaviour.
- [18]
Cross-run contamination in the replay was a property of the harness configuration rather than of the model, since it appeared only when service state persisted and disappeared under isolation or verified reset.
- [19]
Where a route connects runs, the reported number of runs is only an upper bound on the number of independent observations, because connected runs must be grouped into one observation.
Sources
1 independent publisher whose own reporting we read for this story.
- arxiv.orgWhen Is an Agent Evaluation Over?Outcome Finality and Cross-Unit Separation
1 article · August 24, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.
Topics
- Evaluation Reporting TransparencyFollow
- Trial Independence and Statistical UnitsFollow
- Agent Evaluation MethodologyFollow
- Eval Harness Isolation and State ResetFollow
- Benchmark and Demo ValidityFollow