Build1 distinct publisher3 min readUpdated
A new diagnostic benchmark treats the execution layer as something to vary rather than a fixed backdrop. Its conclusion is that agent capability belongs to a model-harness pair.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
The load-bearing part of the paper is the failure taxonomy, not the headline that results move. Harness-Bench reports recurring execution-alignment failures, where plausible reasoning becomes decoupled from tool feedback, workspace state, evidence, or a verifiable output contract [6]. Every item on that list is owned by the layer the authors call the harness: the mechanism that manages context, tools, state, constraints, permissions, tracing and recovery [1]. A base model does not decide whether a tool's error output makes it back into context, whether a stale directory listing gets refreshed before the next edit, or whether an artifact is validated against a contract before the run is marked done. Those are configuration choices, and they are exactly the choices that determine whether good reasoning turns into a correct file on disk.
That explains why the existing evidence does not transfer. Static suites such as MMLU, GSM8K, BIG-bench and HELM measure text-based capability [8]. Executable suites such as SWE-bench, WebArena, OSWorld and Terminal-Bench score complete systems [9], so a good result is unattributable between the model and the scaffold wrapped around it. Workflow suites such as AgentBench, GAIA and Claw-Eval compare model backends under a shared execution setup [10], which makes the resulting model ranking conditional on one scaffold nobody else runs. The authors' framing is that the harness is either abstracted away, conflated with the whole system, or held fixed [4].
The arithmetic of doing it properly is worth noting. 5,194 trajectories across 106 tasks [3][2] works out to roughly 49 runs per task [12]. That is the width of the matrix needed before a configuration-level statement is defensible, and it is a decent explanation for why this has not been done by accident inside product teams: the run count multiplies with every harness variant you take seriously.
Two caveats sit in the design. The comparison holds task environments, budgets and evaluation protocols common while preserving each harness's native execution behaviour [13], which keeps the setups realistic and also means competing configurations differ in more than one variable at a time. And the tasks are sandboxed and offline, reviewed for realism, solvability, oracle-checkability and integrity [2], so whatever rate limits and network flakiness cost you in production is outside this measurement.
The part worth copying regardless of whether the ranking survives is the instrumentation. Each run records final artifacts, execution traces, usage statistics and validator outputs, which the authors say enables analysis beyond final completion [7]. That is the data needed to compute tokens per validated artifact rather than a pass rate, and most internal agent dashboards do not keep validator results next to usage numbers. Code and data are published [11], so the protocol is available to anyone who wants to point it at their own stack.
For now, the supplied text describes the variation qualitatively and does not give per-configuration effect sizes, nor the number of harnesses and models compared [14]. So the direction is stated and the magnitude is not, which is the honest position: the claim on the table is that the choice of model is the half of the decision most teams already have data on, and the cheaper half to change.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
Across 5,194 execution trajectories the authors observe substantial variation in completion, process quality, efficiency and failure behaviour across model-harness pairings.
The authors conclude that agent capability should be reported at the model-harness configuration level rather than attributed to the base model alone.
Harness-Bench evaluates representative harness configurations across multiple model backends under shared task environments, budgets and evaluation protocols, while preserving each harness's native execution behaviour.
Harness-Bench defines the harness as the system layer that manages context, tools, state, constraints, permissions, tracing and recovery, mediating between model outputs and external actions.
The benchmark contains 106 sandboxed offline tasks constructed from practical agent-use patterns and manually reviewed for realism, solvability, oracle-checkability and integrity.
Existing benchmarks either abstract away execution, conflate the harness with the full agent system, or fix the harness when comparing models, leaving the harness itself largely unmeasured.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Single self-reported preprint with concrete methodology but no published numbers
The cluster rests on one arXiv source authored by the benchmark's creators. In its favour: a specific and checkable design (106 manually reviewed sandboxed tasks, fixed budgets/timeouts/evaluators, 5,194 trajectories, per-run artifacts and validator outputs) and a stated code/data release. Against it: no peer review, no independent replication, and the supplied abstract and introduction give no per-configuration effect sizes, no harness identities and no backend counts, so the headline variation claim cannot be inspected.
Author-side release only
The only adoption evidence is the authors' own artifact release and their internal 5,194-trajectory run. No third-party use, integration, citation, download or deployment appears in the supplied material, and the repository's actual state is not independently verified here.
Framing runs modestly ahead of published numbers
The strong reading - that model-only scores lose meaning - is directionally consistent with the paper, but the paper itself is more guarded: it calls the results configuration-level diagnostics rather than causal decompositions, and the supplied text never quantifies how much spread the harness accounts for or which harnesses were compared. A methodological proposal with a self-declared novelty claim and no external validation is being read as a settled measurement result, so the overstatement is moderate rather than severe.
Benchmark authors define the axis on which they claim priority
The work is released under a corporate GitHub organisation (Qihoo360) alongside a dedicated project domain, and the authors claim to be among the first to make the harness a primary evaluation axis. Proposing a new evaluation dimension, selecting the harness configurations tested, and controlling which numbers are published are aligned incentives toward a favourable framing. The supplied text discloses no funding, no conflict statement and no vendor relationships, so the strength of the incentive cannot be bounded further.
Low-moderate: coherent method, unverified results, no external corroboration
Confidence is limited by single-source, pre-peer-review, author-controlled evidence and by the absence of the numbers that would let the central claim be tested. It is not lower because the described protocol is internally consistent, the run scale is specified, artifacts are claimed to be public, and the authors themselves bound their interpretation.
build
Artificial Analysis moves eval onto your data, and turns model choice into procurement1 distinct publisher
build
A 284B model at 25 tokens a second on one 5090, and 192 GiB of DDR5 doing the quiet part1 distinct publisher
security
Two Artifactory flaws poisoned metadata, not artifacts, and that was enough to break a shared cache1 distinct publisher
build
Uniform INT4 beat NVFP4 in 85 of 90 real gradient tests, and the rotation barely moved them1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 23, 2026