Build1 distinct publisher3 min readUpdated
Locally Uncensored put four runs of one bugfix on the same model at the same prices. Its own 2.6.5 agent cost 4395 credits, about 3.4 times its successor and double opencode's average.
The Engineer · Build desk
Compiled by The EngineerSomething wrong?How this is made
The arithmetic that makes this table move is in the payload, not the plan. opencode's cheapest run needed 98,789 tokens across eight requests [16], which works out to about 12,349 tokens per request [3]. The vendor's own note that six opencode steps at that weight roughly equal its total consumption for the whole task [9] puts its sixteen requests somewhere near 4,500 tokens each [4]. Same billing rate: credits per prompt token land within a few percent across all four runs, and the vendor's is marginally the highest of the set [6]. Nobody got a discount. The invoice moved on volume alone [6].
Two things set that volume, and the vendor names both: a tool catalogue of 21,188 bytes against 7,703, roughly triple the standing overhead and charged on every call whether the model touches those tools or not [7], and a transcript that keeps resending stale tool output at full length unless something trims it [8]. That is why the step-count intuition inverts here. opencode used eight to eleven requests against the vendor's sixteen [5], so scoring on turns hands opencode the win while the bill goes the other way. At opencode's highest observed turn count its average run still costs about 196 credits per request, against 81 for the 2.6.6 agent [5].
The row worth reading twice is 4395, which is not a constructed strawman but what the vendor shipped one release earlier as 2.6.5 [4]. That is about 3.4 times its own successor on the identical task [1] and roughly double opencode's three-run average [6]. The vendor says the 2.6.5-to-2.6.6 work targeted exactly those two items and cut credit consumption 78.6 percent across its tool-driven runs, 80.4 percent on the longest [10]. On this particular bugfix the cut is 70.5 percent [2], noticeably less than the headline figure. A vendor building a chart to flatter itself would not usually publish the version of the number that is worse.
The limits are stated rather than buried, and they are real. opencode ran three times, the vendor once, and the spread inside opencode alone is 45 percent, from 1679 to 2433 credits, while the vendor's own variance is unknown [12]. opencode ran as it ships, and a trimmed tool set would land elsewhere [13]. The scenario is one tiny repo and a one-line bug, with nothing said about multi-file refactors or hour-long sessions [11]. opencode is free software, so what is measured is the model bill and not a licence [14]. Every run in the table produced correct, committed work under criteria fixed before the runs [2][15].
What survives all that is narrower than 40 percent and more useful. On short, well-scoped tasks, opencode can hardly land below the vendor because the fixed per-request overhead sets a floor [17]. That mechanism does not care whose agent it is, which is the point of the 2.6.5 row: the same measurement pointed at your own shipped loop will find the same standing block. Anyone running an agent in production can price this without a benchmark harness by measuring the size of the block sent on every call and whether anything at all trims the transcript behind it.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
opencode averaged 2157 credits over three runs; the vendor's 2.6.6 agent finished the identical task for 1298 credits, about 40 percent less, and the cheapest opencode run came in 29 percent above the vendor's number.
The vendor ran the same bugfix, same model, same API, same prices and a byte-identical prompt once through opencode and once through the coding agent inside Locally Uncensored.
Success was defined before the runs: npm test passes, exactly one commit with the required message, only add.js changed, clean working tree at the end. All four runs cleared that bar and nothing failed.
The 4395-credit row is what the vendor shipped in its previous release (2.6.5) and is the most expensive entry in the table, not a strawman built for the article.
opencode used eight to eleven requests to complete the task; the vendor's agent used sixteen.
The billing rate was the same across all four runs: credits per prompt token land within a few percent of each other and the vendor's is marginally the highest of the set, so the entire difference in the invoice is token volume, not token price.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Granular numbers, single self-published source
The source is unusually specific for vendor content: pre-registered pass criteria, per-run credit totals, prompt-token counts, credits-per-prompt-token rates and tool-catalogue byte sizes, plus a limitations section that names the sample and configuration weaknesses. But it is one vendor-authored post with no independent replication, the vendor's own agent ran once, the task is a one-line bugfix in a tiny repo, and the raw counts behind the internal 78.6 percent claim are not shown.
No uptake evidence supplied
The supplied material contains a benchmark and a release description only. There are no downloads, install counts, user or customer disclosures, production deployments, or third-party usage reports for either the vendor's agent or opencode, so real-world adoption cannot be scored without inference.
Headline generalises past a one-task sample
The '40 percent fewer credits' framing is presented as a comparative verdict but rests on one tiny-repo bugfix, untuned opencode defaults, and a 3-to-1 run asymmetry, and the vendor's floor generalisation for 'short, well scoped tasks' outruns what a single scenario can establish. The overstatement is modest rather than severe because the same article publishes the counter-evidence itself: its own prior release is the most expensive row, opencode took fewer turns and never failed, opencode is free, and cost is explicitly separated from quality.
Vendor self-benchmark with product funnel
The measurement, the framing and the publication all come from the party whose product wins the headline comparison, and the piece routes readers to the vendor's agent, its full writeup and its hosted LU Labs Cloud model picker. Mitigating factors are visible but do not remove the incentive: the vendor publishes raw counts, discloses that opencode is free software and took fewer turns, and shows its own previous release as the worst performer.
Low-moderate
Confidence is limited by structure rather than by vagueness: a single self-interested publisher, one scenario, n=1 on the vendor's side, and no independent verification path exercised in the supplied material. The internally consistent and checkable token arithmetic, the pre-registered success criteria and the explicit limitations list keep it above the floor, and the mechanism (fixed per-request overhead plus transcript resend) is more credible than the specific credit ratios.
build
Qwen 3.8's Apache-licensed 27B is the one you can actually own, and its KV cache is why1 distinct publisher
build
Your agent needs the API call, not the API key1 distinct publisher
build
Waku 0.1.0 bets the product is the control plane, not another coding agent1 distinct publisher
build
Cloudflare moves durable execution under the harness, and the platform starts choosing it1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 22, 2026