Build1 distinct publisher3 min readPublished
The spread traces to two numbers you can read off your own logs: the tokens a harness spends before any work starts, and how many turns it takes. Together they predicted total token use with an R-squared of 0.99.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
Divide the endpoints and the number gets bigger. 292,000 tokens per solved task for OpenClaw against roughly 3,500 for Aider in architect mode is a factor of about 83 [6][1], not the 70-fold in the writeup's own headline [23]. Round numbers travel better than division does. Either figure sits far outside anything a buyer wins by negotiating price per million tokens.
The floor recurs because of how serving works. Every inference request needs its context, and the relevant history is either supplied again or reconstructed by the serving system, so the provider processes large overlapping blocks of text on each turn, including the harness's system prompt and its tool descriptions [2]. The startup tax is that baggage measured once: the system prompt, the tool descriptions and the environment setup, around 700 tokens for Aider in architect mode against around 26,000 for OpenClaw [7]. The June author's regression multiplies it by turn count and predicts tokens per solved task with an R-squared of 0.99 across both models [9].
You can run that regression backwards. Divide 292,000 by 26,000 and you get about eleven [4]. So on the tasks OpenClaw solved, roughly eleven resends of the same scaffolding account for most of the token spend, and the actual task content is what is left over. The fifteen-turn illustration in the post, 390,000 input tokens of scaffolding [8], is above the measured per-task total, which is a useful reminder that it was an illustration and not a measurement.
Dollars did not spread as far as tokens. Composio's eight-harness comparison put cost per successful task between $0.028 for Pi Agent and $0.195 for Claude Code, a range of about 7x [16][3]. Different harnesses, and different work: thirty workflows across Airtable, Gmail, Google Calendar, Google Sheets, GitHub, Slack and PostHog, each under a 900-second ceiling [13]. DeepAgents matched Claude Code's pass rate exactly at a quarter of the cost per success, so around five cents [17][5]. Composio graded with a programmatic verifier rather than an LLM judge, on isolated fixtures seeded with decoys and near-identical keys [14], and it disclosed that Pi ran a different reasoning setting across two model providers and that Prime Agent produced only 24 gradable runs out of 30 [18]. That is careful benchmarking, and it is also why the result cannot be read as single-variable.
For the token ordering to transfer to your invoice, your turn distribution has to resemble theirs. The June suite was twelve Python tasks with one pass per harness, task and model combination, so it carries no variance estimates while agent runs are stochastic [3][10]. Artificial Analysis is the better instrument for ranking live pairings, at 326 tasks with pass rates averaged over three attempts, plus a controlled swap that holds Claude Opus 4.7 fixed across Claude Code, Cursor CLI and Opencode [19][21][24]. What transfers without any benchmark is the method: measure your harness's prompt floor, then your median turns to done [11].
Ranked by verification strength, evidence, and original report placement.
The June author noted that a harness carrying a 26,000-token floor through fifteen turns spends roughly 390,000 input tokens on scaffolding alone.
The June post described the scaffolding overhead as a 40x difference that would be tolerable if paid once, but the resend pattern makes it otherwise.
Three recent benchmarking efforts suggest the harness, the software that steers a model through tasks, may matter as much as the model when pricing an AI coding agent.
Each inference request needs context, and the relevant history must either be supplied again or reconstructed by the serving system, so the provider processes large overlapping blocks of text on every turn, including the harness's system prompt and tool descriptions.
In June, an independent benchmark compared 12 configurations across two models on the same 12 Python tasks.
The June benchmark ran Aider, Claude Code, Codex, Goose, Hermes, Kilo, Kimi Code, Nanobot, OpenClaw, Opencode and Qwen Code, counting Aider's architect mode separately, with all twelve configurations going through OpenRouter on the same tasks so each harness used the same API and model.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 27, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
invest
DeepSeek V4 Flash costs a tenth as much and passes 53.8% of agent tasks1 distinct publisher
build
Superpowers makes spec-driven work a precondition, then ships it to twelve harnesses1 distinct publisher
security
An agent guard that runs on your laptop, and cannot tell you whether anyone keeps it on1 distinct publisher
build
Seven harnesses took the transcript. None of them took the approvals.1 distinct publisher
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Convergent benchmark evidence, single-publisher and partly single-pass
Three independent measurement efforts point the same way, methods are described concretely (shared OpenRouter routing, programmatic verifier, 326 tasks averaged over three attempts), and the mechanism has a quantified fit (R-squared 0.99). But the whole cluster rests on one publisher's account, the token-spread benchmark ran a single pass with no variance estimates, the vendor comparison has disclosed control breaks, and the source's own headline figure (70x) conflicts with its reported endpoints (~83x).
Widely benchmarked tooling, no production usage disclosed
Adoption evidence here is confined to test harnesses being exercised in benchmarks: twelve configurations run through OpenRouter, eight harnesses across thirty enterprise-style workflows, and a continuously tracked index of harness-model pairings. That shows the tools are real, runnable and comparable at benchmark scale, but the sources disclose no deployment counts, customer usage, revenue or production spend for any harness, so real-world uptake is unmeasured.
Mechanism modestly oversold by the framing, not by the numbers
The underlying measurements are reported carefully and the article volunteers its own limits (single pass, stochastic runs, vendor caveats, index preferred for production comparison). The overstatement is at the framing layer: a headline spread figure (70-fold) that does not match the body's endpoints (~83x), extrapolation from twelve Python tasks and thirty fixtures toward general agent pricing, and no accompanying task-quality comparison to show cheap harnesses solve equivalent work.
One vendor-run comparison, disclosed; other instruments more arm's length
Composio published the August comparison it is cited for, which is a self-interest exposure, though it disclosed the control breaks (Pi's differing reasoning setting, Prime Agent's 24 gradable runs of 30) transparently. The June benchmark is described as independent and rerunnable at no cost on a free Nemotron 3 Ultra tier, and Artificial Analysis is a third-party index — both reduce distortion risk. No pricing, sponsorship or commercial relationships between the publisher and any named harness vendor are disclosed either way, so residual uncertainty stays moderate.
Moderate: mechanism credible, magnitudes provisional
The causal story — resent scaffolding floors times turn count driving token spend — is coherent, quantified and reproduced in ordering across two unrelated models, so confidence in direction is fair. Confidence in the specific magnitudes is lower: one publisher, one single-pass token benchmark, a vendor comparison with disclosed caveats, an internal headline/body discrepancy, and no independent replication or vendor rebuttal in the cluster.