Build1 distinct publisher3 min readPublished
The Accio team's verifier ignores what the agent says it did and inspects the containers it leaves behind. That design is the news; the two-pass open-weight lead Alibaba reports from it sits inside its own harness spread.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
The graders sit outside the agent's own view of the system. Task graders, rubrics, private seeds and mock implementations stay outside the tree the agent can see [5]. The agent gets a business request, a workspace and access to local software replicas, and every run begins in a fresh container [4]. When it stops, a host-side verifier reads the final state of the mock services and any files produced, and the task passes only if every required check passes [6]. The familiar failure mode, where a model narrates a plausible sequence of actions and reports success while the database or storefront is untouched, scores zero here [7]. Lian and Xie also contribute to Business Arena, which extends the same state-based thesis to a simulated cross-border shop over a longer operating period [26].
For the scores to mean anything on your stack, your stack has to resemble the 14 offline replicas the suite ships, covering product publishing, freight booking, storefront administration and payment operations [8]. Then there is the mix: 53 of the 107 tasks are command line, 28 browser, 16 file and document, 10 API or MCP [3]. Command-line work is 49.5% of the suite [25], so an aggregate number is mostly a statement about shell competence. If your agents drive a browser session against a live storefront, that aggregate is answering a question you did not ask.
Now the ranking. Qwen3.8-Max's 156 passes across 321 harness-task combinations work out to 48.6% [10]; DeepSeek V4 Pro's 154 come to 48.0% [12]. The gap sits entirely inside one harness: ties in Pi and OpenClaw, then two tasks ahead in Accio [13]. That is six tenths of a percentage point [23]. The spread inside Qwen's own row is nine tasks between Accio and OpenClaw [18], four and a half times the margin that decides the open-weight ordering [24]. Claude Opus 5, which moved six tasks between its weakest and strongest harness [19], finished 35 passes ahead of Qwen regardless [15], at 59.5% and 10.9 points clear [14].
The code is Apache 2.0 and the task data CC BY 4.0, with commercial use permitted [9], so nothing stops a rerun. What runtimewire reports as missing is an immutable public bundle at task level [22], and its read of the tables is that they come from Alibaba's own benchmark and should be treated as a published result rather than an independent audit [21]. The practical effect is that nobody outside the team can name the two Accio tasks carrying the strongest-overall claim. Checking them means standing up 321 runs of your own, for a result that separates the harnesses by six tenths of a percentage point.
One integration detail is worth flagging while hunting for the source of harness spread. The repository's Qwen documentation notes the model returns its reasoning trace in a separate response field [20]. A harness that reads that field and a harness that ignores it are not running the same agent.
The ceiling is the figure I would take to a review. The best result anywhere in the tables still left 41 of 107 tasks unfinished [16], and the source's own conclusion is that a 61.7% pass rate over listings, freight, payments and customer records demands supervision and transaction-level checks [17].
Ranked by verification strength, evidence, and original report placement.
Alibaba's official Qwen account described Qwen3.8-Max's Commerce Agent Bench results as the strongest overall performance among open-weight models.
The figures come from Alibaba's own benchmark, so the ranking is best read as a published result rather than an independent audit, according to runtimewire.
Runtimewire says Alibaba's self-reported two-pass open-weight margin, harness variance and the lack of immutable public task-level bundles make auditability as important as leaderboard order.
Yukun Lian, Sicong Xie and ten colleagues on Alibaba International's Accio team built Commerce Agent Bench to test whether AI agents can finish commercial work inside stateful replicas of business software.
Commerce Agent Bench contains 107 tasks: 53 command-line tasks, 28 browser tasks, 16 file and document tasks, and 10 API or MCP tasks, covering work such as publishing products, booking freight, editing storefronts, analyzing suppliers, researching the public web and producing spreadsheets.
Each task starts in a fresh container; the agent receives a business request, a workspace and access to local software replicas.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 31, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
product
A 27B laptop model scores like a rented one, and thinks three times as hard to do it1 distinct publisher
science
Frontier scores at $2/$6: Grok 4.6 ties GPT-5.6 on one evaluator's index for a fifth the output price1 distinct publisher
invest
The cheap-token trade is closing: DeepSeek's 12x price rise resets everyone's AI cost model1 distinct publisher
product
Alibaba is selling a working games studio to pay for Qwen1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One outlet, reading the vendor's own repository
Every figure in this story — 107 tasks, 14 replicas, 156 passes, 48.6% — originates in the repository Alibaba's Accio team publishes about its own model, relayed by a single publication. The internal arithmetic is checkable and consistent, and runtimewire quotes the methodology against itself, but the materials concede the raw per-task bundles behind the board are neither in Git nor archived with checksums, so no outsider can reconstruct a single row.
Published and licensed; only the authors have run it
A benchmark is adopted when other people score against it. Here the only runs on record are the authors' own, across three harnesses — one of which the materials themselves label reference-only and unreproducible from a checkout. Apache 2.0 code and CC BY 4.0 tasks make reuse legally easy and OpenClaw is described as rerunnable, but no third party has published a number, and the invitation to contribute task domains is still an invitation.
The ranking is oversold; the design is undersold
'Strongest overall performance among open-weight models' is carrying two passes out of 321, sourced entirely to the harness that cannot be rerun, while the same tables show Qwen swinging nine tasks between its own scaffolds and a closed model finishing 35 passes ahead. The gap is not fabrication — the numbers are real and runtimewire reports them fairly — it is a claim of leadership resting on noise the benchmark's own spread dwarfs. The verifier that ignores the transcript and reads the container deserves more attention than the row it produced.
Author, grader, harness operator and entrant are the same team
Alibaba's Accio team wrote the tasks, wrote the hidden graders, built and ran the harnesses, and entered its own flagship model — then its corporate Qwen account announced the win. Add the disclosed programme of private pre-release evaluations for other people's checkpoints and the benchmark becomes a business surface as well as a measurement, which is precisely when the missing checksummed task bundles start to matter.
Solid on what was reported, thin on who reported it
We are fairly sure the benchmark exists as described and that the tables say what runtimewire says they say; the per-harness rows are specific, the aggregates reconcile, and the caveats are the publisher's own rather than ones we had to supply. What keeps this in the middle is structural: one publication, one primary source with a stake in the result, and no rerun by anyone. A second independent score on OpenClaw would move this more than any further reading of the repository.