Build1 distinct publisher3 min readPublished
A preprint proves a blind replay agent's expected score equals the source model's pass@k, which means static computer-use leaderboards have been grading environment determinism.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
A program with no perception cannot adapt to anything. It scores only because there is nothing to adapt to. On a static benchmark, every task starts from one fixed state, and the paper's argument is that this alone lets a memorised action sequence pass for visual reasoning [3]. Remark 1 in the preprint turns the observation into an identity: in a deterministic environment, the expected success rate of the replay agent is exactly the source agent's pass@k [2]. The pass@k figure printed beside a frontier model on such a suite is therefore also the score available to a 1MB file that never reads the screen [1].
Follow the arithmetic one step further. If the source model is credited with a single attempt and the replay is credited with a recorded success, the margin between them is pass@k minus pass@1 for that same model [1]. Nothing about the script produces that margin. It is a readout of how inconsistently the source model repeats itself, which is why the small agent can beat the large one it was recorded from.
That is a metric transplanted out of its habitat. pass@k was built for stateless code generation, where a discarded sample costs nothing and the next one begins from a clean slate; the authors argue it does not survive the move to stateful UI interaction [4]. The environment side compounds it. Suites tied to live services drift and stop being reproducible [5]. Parameterised configuration spaces without automated integrity checks leave some fraction of tasks impossible, incoherent, or already solved at the start, which quietly corrupts the aggregate [6]. Verification by LLM judge or screenshot comparison carries its own biases and can be attacked directly [7].
The audit result is the part that should worry anyone who cites these numbers: the paper reviews the major CUA benchmarks and reports that not one of them satisfies the full set of requirements [8], a zero out of ten across the works it cites [3]. It also points to concurrent audits, citing Wang et al. and Stein et al., finding that benchmarks can be driven to near-perfect scores without solving any task, and that this is already common practice [9][10].
The proposed replacement shows what a defensible harness costs. DigiWorld is fifteen sandboxed mobile applications carrying more than 3.2 million verified configurations [11], which averages roughly 213,000 per application [2], and the accompanying statistics pair Wilson score intervals with a hierarchical bootstrap so that nested trials are not counted as independent coin flips [12]. Two things follow. Building that is a different discipline from writing a task list, and it is being offered by the same authors whose diagnosis creates demand for it, with "verified" resting on their own integrity checker [13]. Meanwhile, the historical scores cannot be repaired after the fact. Without knowing whether a given environment re-randomised between trials, there is no way to tell how much of any published CUA number was the agent.
Ranked by verification strength, evidence, and original report placement.
The paper proves (Remark 1) that the expected success rate of the replay agent is exactly equal to the source agent's pass@k in deterministic environments, revealing that pass@k on static benchmarks measures memorization capacity rather than agent capability.
Among existing major CUA benchmarks the paper inspects, none satisfies all of its stated design principles: realistic, sandboxed and varied enough to prevent memorization, with configurations verified as solvable and completion checked against internal application state.
The paper states that concurrent systematic audits have shown major benchmarks can be exploited to achieve near-perfect scores without solving any tasks (Wang et al., 2026).
The paper states that such benchmark exploitation is already widespread (Stein et al., 2026).
DigiWorld is a benchmark of 15 realistic sandboxed mobile applications able to evaluate agents in over 3.2 million verified unique configurations.
The same paper that diagnoses the benchmark failures proposes the PRISM design principles (privileged verification, realistic environments, integrity-checked configurations, sandboxed execution, multifactorial variability) and instantiates them in its own benchmark, DigiWorld.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One primary preprint: a formal result plus authors' own experiments, no independent replication
The central claim is backed by a stated proof (Remark 1) and an empirical demonstration on popular deterministic benchmarks, which is stronger than assertion alone, and the environment-design failure modes are itemized concretely. But the cluster holds a single unreviewed source authored by the parties proposing the remedy; the supplied text omits k values, pass@1 baselines, per-benchmark scores, and the per-axis multiplicities needed to recompute 3.2M configurations, and the corroborating exploitation audits appear only as second-hand 2026 citations.
No third-party adoption evidence in the cluster
The only observable event is the authors' own preprint release introducing PRISM and DigiWorld. The supplied material discloses no downstream users, no other lab or vendor adopting the principles or the aggregation framework, no leaderboard integration, and no availability, license, or usage figures for DigiWorld, so adoption cannot be scored without guessing.
Slightly overstated: tight formal core, self-scored comparative verdict
The framing that a screen-blind 1MB script beats frontier models is arresting but conditional on a specific accounting — replay credited with a recorded success against a source scored over k attempts — and the supplied text does not disclose those parameters, so the gap between headline and shown numbers is real. The sweeping verdict that none of ten prior benchmarks meets the bar is graded against a rubric the authors authored and their own benchmark satisfies. The mathematical remark itself is modest and specific, which keeps the gap small rather than large.
Author incentive: the paper that fails all incumbents ships the replacement
The same authors define the design rubric, judge ten incumbent benchmarks as failing it, and publish DigiWorld plus PRISM as the instantiation that satisfies it, alongside an aggregation framework they report outperforms widespread practice. That is a clear structural incentive to maximize the severity of incumbent flaws and the novelty of the remedy. No funding, vendor relationship, or commercial interest is disclosed in the supplied text, so the assessment rests on the visible authorship structure only.
Moderate-low: internally coherent but unreplicated and single-publisher
Confidence is held down by the cluster's structure: one publisher, one self-interested preprint, no peer review, no independent replication of the replay result, and no adoption signal to triangulate. It is held up by the fact that the load-bearing claim is a stated formal identity with a clear mechanism, and by the specificity of the disclosed benchmark and aggregation details, which are checkable once the artifact is examined.
build
A GPU SQL Engine Lost to One CPU Thread Because a Dispatcher Constant Was 128x Too Small1 distinct publisher
science
GJ 523b gives 'Mega-Earth' a number: 23 Earth masses inside 2.5 Earth radii1 distinct publisher
build
A refactoring benchmark stops the best agent at 41.2%, and the tests are the story1 distinct publisher
build
Multi-agent LLM gains largely vanish once the thinking-token budget is held constant1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 25, 2026