Skip to content

BuildNot yet confirmed elsewhere1 publisher2 min readPublished

Ten agent eval protocols say when a run stops. Fewer say whether the result is settled.

A replay that held the agent's actions fixed produced different labels before and after delayed operations resolved, and one late write moved the following run's score.

The Engineer · Build desk

How we use AISend a correction

What happened

  • An arxiv paper argues an end-of-run score counts as a final result only if outcome finality and cross-unit separation hold, and stopping a run establishes neither.
  • In a controlled replay, the label read at the endpoint disagreed with the label read after the operation resolved in every delayed case.
  • A write applied after the endpoint changed the next run's score whenever service state carried over between runs.
  • Across ten public protocols, the inconsistently documented parts were unfinished operations and the evidence that runs are separate trials.

Why it matters

  • constraint Once a route between runs is admitted, the run count becomes a ceiling on independent observations rather than the observation count, so error bars and win margins computed from trial counts are...
  • exposure Anyone reusing a published number from a shared-service harness inherits a contamination the number has no field to declare, including procurement teams and safety reviewers who never touch the...
  • decision Evaluators now face a three-way call on runs with operations still in flight: re-score after resolution, bound the effect, or publish the trial as uncertain instead of as a pass or a fail.
  • contradiction The paper leans on incident reports to show the boundary really is crossed in production, then rules out using them to size the problem, which supports new record-keeping but not discounting any...

The replay holds the agent's actions fixed and varies only the accounting [1]. That is what makes it a mechanism demonstration rather than a benchmark result: if the actions did not change, the divergence between the two labels cannot be charged to the model behaving differently [17]. Same trajectory, two labels, and the one that gets published depends on when the evaluator looked.

The second half of that finding is the part an infrastructure team can act on. The carryover into the next run appeared when service state persisted and did not appear after isolation or verified reset [c5b], which places the defect in the harness configuration rather than in the system under test [18].

The two conditions also do not come as a pair [6]. Polling until the write lands settles whether the write succeeded, and leaves the written state sitting where the next run can read it. A fresh container prevents that read, and tells you nothing about whether the operation the score refers to ever finished [6]. Each of the two obvious fixes buys exactly one condition.

Which is roughly why the protocol review found what it found. The endpoint is defined by the harness: the evaluator stops asking for actions after the model reports completion, or on a success condition, a timeout, or an action limit [9][10]. All ten protocols document that boundary and what gets scored [c6a], because all of it is in code the evaluator owns. The status of a pending write, a leftover credential, or an account created mid-run sits in a service the evaluator does not own [9]. So counting a stopped run as one trial [4] is an assertion about the environment, offered in place of evidence about it [15].

The two field reports the authors cite map cleanly onto the two conditions. The UK AI Security Institute's concurrent samples reusing accounts and artifacts left by other samples is separation failing [11]. The OpenAI and Hugging Face accounts of evaluation activity that used external services and continued past the original runtime is persistence past the endpoint, the case where the scored outcome was still moving [12]. The proposed open-effects record does not repair either one. It lists what may still be live after the endpoint, its status, and whether it could change the score or reach another run [16]. That is a disclosure format, and its whole value is that it gives a harness somewhere to write down what it does not know.

What to watch

  • Whether any of the ten reviewed protocols adopts an open-effects record, or starts publishing pending-operation status alongside scores.
  • Whether anyone measures how often endpoint and terminal labels diverge outside a controlled replay, which would put a rate on the mechanism.
  • Whether leaderboards begin asking for evidence of isolation or verified reset before accepting a reported trial count as a sample size.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence46
Adoption
Insufficient
Hype gap−8
Incentives34
Confidence50
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    In a controlled replay where the agent's actions were held fixed, the endpoint label and the terminal label differ for every delayed operation.

  2. [2]

    In a review of ten public protocols, all protocols identify when a run stops and what is scored.

  3. [3]

    In the same review of ten public protocols, unfinished operations and the evidence for treating runs as separate trials are documented less consistently.

Sources

1 independent publisher whose own reporting we read for this story.

  1. arxiv.org

    1 article · August 24, 2026

    When Is an Agent Evaluation Over?Outcome Finality and Cross-Unit Separation

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Topics

Entities

Loading related stories