Skip to content

Build1 publisherNot yet confirmed elsewhere3 min readPublished

Harness-Bench makes the scaffold a measured variable, and model-only scores lose meaning

A new diagnostic benchmark treats the execution layer as something to vary rather than a fixed backdrop. Its conclusion is that agent capability belongs to a model-harness pair.

The Engineer · Build desk

How we use AISend a correction

Illustration accompanying Harness-Bench makes the scaffold a measured variable, and model-only scores lose meaning
Generated illustration

What happened

  • It runs 106 sandboxed offline tasks built from practical agent-use patterns and manually reviewed for realism, solvability and oracle-checkability.
  • Across 5,194 execution trajectories the authors report substantial variation in completion, process quality, efficiency and failure behaviour by model-harness pairing.
  • Their recommendation is that capability be reported per model-harness configuration instead of credited to the base model.
  • The code and task data are published, along with a project site.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint Because the published suites either fuse the scaffold into the score or freeze it, no existing agent number can be carried across to a stack that wires context and permissions differently.
  • cost Treating the harness as a variable multiplies run counts rather than adding to them, and the compute bill for a defensible comparison falls on whichever team wants a number that transfers.
  • capability Keeping usage statistics next to validator outputs makes cost per verified artifact computable, which is the figure a pass rate quietly hides from anyone budgeting an agent rollout.
  • precedent If configuration-level reporting catches on, a vendor claim about agent performance becomes unreadable without the scaffold disclosure attached to it.

The load-bearing part of the paper is the failure taxonomy, not the headline that results move. Harness-Bench reports recurring execution-alignment failures, where plausible reasoning becomes decoupled from tool feedback, workspace state, evidence, or a verifiable output contract [7]. Every item on that list is owned by the layer the authors call the harness: the mechanism that manages context, tools, state, constraints, permissions, tracing and recovery [4]. A base model does not decide whether a tool's error output makes it back into context, whether a stale directory listing gets refreshed before the next edit, or whether an artifact is validated against a contract before the run is marked done. Those are configuration choices, and they are exactly the choices that determine whether good reasoning turns into a correct file on disk.

That explains why the existing evidence does not transfer. Static suites such as MMLU, GSM8K, BIG-bench and HELM measure text-based capability [9]. Executable suites such as SWE-bench, WebArena, OSWorld and Terminal-Bench score complete systems [10], so a good result is unattributable between the model and the scaffold wrapped around it. Workflow suites such as AgentBench, GAIA and Claw-Eval compare model backends under a shared execution setup [11], which makes the resulting model ranking conditional on one scaffold nobody else runs. The authors' framing is that the harness is either abstracted away, conflated with the whole system, or held fixed [6].

The arithmetic of doing it properly is worth noting. 5,194 trajectories across 106 tasks [1][5] works out to roughly 49 runs per task [14]. That is the width of the matrix needed before a configuration-level statement is defensible, and it is a decent explanation for why this has not been done by accident inside product teams: the run count multiplies with every harness variant you take seriously.

Two caveats sit in the design. The comparison holds task environments, budgets and evaluation protocols common while preserving each harness's native execution behaviour [3], which keeps the setups realistic and also means competing configurations differ in more than one variable at a time. And the tasks are sandboxed and offline, reviewed for realism, solvability, oracle-checkability and integrity [5], so whatever rate limits and network flakiness cost you in production is outside this measurement.

The part worth copying regardless of whether the ranking survives is the instrumentation. Each run records final artifacts, execution traces, usage statistics and validator outputs, which the authors say enables analysis beyond final completion [8]. That is the data needed to compute tokens per validated artifact rather than a pass rate, and most internal agent dashboards do not keep validator results next to usage numbers. Code and data are published [12], so the protocol is available to anyone who wants to point it at their own stack.

For now, the supplied text describes the variation qualitatively and does not give per-configuration effect sizes, nor the number of harnesses and models compared [13]. So the direction is stated and the magnitude is not, which is the honest position: the claim on the table is that the choice of model is the half of the decision most teams already have data on, and the cheaper half to change.

What to watch

  • Whether the full paper publishes per-configuration effect sizes and whether harness ordering holds across model backends.
  • Whether anyone reproduces the comparison on a production stack using the released code and tasks.
  • Whether vendors start attaching harness configuration details to reported agent benchmark results.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories