Skip to content

Build3 publishers2 min readPublished

Two-thirds of failed agent runs in ThinkingBox exited cleanly with the backend still wrong

Microsoft and Hugging Face's ThinkingBox found that 67.24% of 79,853 failed agent runs ended cleanly, with no final tool error. Those failures showed up only when executable checks read the records each run left in the backend.

The Engineer · Build desk

Illustration accompanying Two-thirds of failed agent runs in ThinkingBox exited cleanly with the backend still wrong
Generated illustration

What happened

  • In the worked example, the agent made nine tool calls on a late-delivery complaint and correctly found that the customer's account segment did not qualify for compensation.
  • Among failed runs, the checks found wrong field values in 77.61%, unintended extra effects in 43.30% and missing required effects in 25.36%, with the categories overlapping.
  • Claude Opus 5.5 led the single-attempt pass@1 ranking at 67.16%, two-thirds of a point above Claude Opus 5.
  • Claude Opus 4.6 scored 68.62% on the retail tasks and 8.30% on the auto-insurance tasks.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • exposure Failures that exit cleanly do not trip error-based monitoring, so in production they stay in customer records until someone inspects the records themselves.
  • cost Grading for consistency this way costs 10,140 runs per model, each from a freshly reset backend, against 507 for a single pass over the tasks.
  • decision On the authors' task mix, a team that needs open weights gives up less than a point of pass@1 by choosing Kimi-K3 over GPT-6-Astra.
  • constraint State grading only works where the correct end state can be written in advance as a field value, such as hold for a ticket blocked on an open carrier exception.

The check that fails in the worked example is one field. The ticket's status is solved, where the task requires hold [6]. The carrier exception was still open, so the correct end state was on hold, pending resolution [5]. The agent's last message read: "Since your query is resolved, is there anything I may assist you with?" [7] It is a courteous sentence to send someone whose $745 appliance is fifteen days late and still stuck at a Nashville distribution center [3]. According to the authors, she also never got an answer to what she asked [8].

Grading the transcript cannot catch this. A grader that reads the tool calls would find nothing malformed. One that checks whether the agent wrote to the database would find the write [19]. Only the value in the status field shows the error [6].

ThinkingBox gets to that value by controlling the environment. The agent works against isolated MCP tool sessions. When the run ends, the grader reads the terminal backend state and the side effects left behind [1]. Each of the 507 workflows runs 20 times, and every run starts from an identical clean backend [9]. One of the three reported numbers, observed 20/20, is the literal count of tasks a model passed on all 20 runs, with no estimator and no smoothing [10]. I think a literal count is the right design for a reliability figure, because nobody needs an error model to read it.

The ablation shows how large the gap is. Across 121,680 valid trials on 12 models, 79,853 runs failed the executable checks [11]. The failure rate is about 65.6% [2]. The trial count equals 507 tasks times 20 runs times 12 models [1]. Roughly 53,700 of the failed runs changed state, terminated cleanly and reported no final tool error [3][12]. That is about 44% of every valid trial [4].

The authors' summary of the method runs to three sentences: "A trajectory is a claim. Database state is the evidence. Repetition is the trust test." [14]

The pass@1 table needs more care than the state checks. It describes the authors' mix of 507 workflows. For an overall score to carry over to a team's own agent, that team's tasks would need a similar domain mix and tools exposed the same way, through MCP [1]. The domain breakdown argues against borrowing the overall figure. One model's retail and auto-insurance scores sit about 60 points apart [5].

The harness is public. The benchmark runs through OpenEnv. The worked example comes from task test_case_ST003_006 in sandbox_external_retail_group1.py, and the full trace is in Appendix D.4 of the paper [18].

What to watch

  • The observed 20/20 counts per model, showing how many of the 507 tasks each model completes correctly on every run.
  • Independent OpenEnv runs of ThinkingBox on other backends, to test whether the clean-exit failure share holds outside the authors' workflows.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories