Build1 distinct publisher3 min readPublished
A dev.to harness put five coding tasks through four APIs and every run passed its verifier on the first attempt, so the only thing left to compare is clock time, where one trial per cell sets a fragile order.
The Engineer · Build desk
invest
One Anthropic order, six times the price: Fractile's $6.5B mark arrives two years before its chips3 distinct publishers
build
WRITER's new flagship is a post-train of Z.ai's GLM-5.2, and that is the story1 distinct publisher
invest
GLM-5.3 Buys Buyers Time: Z.ai's Coding Model Cuts Tokens, Not the Closed-Model Lead1 distinct publisher
leadership
Ord's generation-time argument makes runaway AI unlikely, not just slower1 distinct publisher
Compiled by The EngineerSomething wrong?How this is made
Latency here comes down to time-to-first-token, not decode rate. NVIDIA NIM averaged 27 seconds of wall time with a 23-second mean time-to-first-token [8], so roughly 85% of the clock elapsed before the first token arrived [2]. That is thinking-token budget and queueing, not tokens per second. Groq makes the same case from the other end: 511 tok/s sustained during generation, and still a 5.8-second mean TTFT, because one task opened a long thinking block [7]. Buy on tok/s and you are paying for the part of the pipeline that generation speed never bottlenecked.
The spread also needs one number lifted out of it. NVIDIA's cli-flag run took 61.7 seconds against OpenRouter's 2.7, which the author reports as a 23x gap on the same prompt [9]. Against a 27-second five-task mean [8], NVIDIA's other four tasks average (135 - 61.7)/4, or 18.3 seconds [3]; OpenRouter's other four average 2.95 seconds [4]. Drop the worst cell on each side and the ratio is 6.2x [5] rather than the 9.3x the means give [1]. Both figures are honest, just answers to different questions, and which you want depends on whether multi-file edits are your common case or your tail. The paid control sat in the middle, with Claude through the Cline API single-billed aggregator averaging 7.4 seconds, about 2.5x the free OpenRouter route [6].
There is one trial per cell, 20 cells in total [4]. With only one trial per cell, 61.7 seconds is simply the sole observation recorded there, and outlier is the wrong label for a one-off sample. Worth saying plainly, because the correctness half of the harness is stricter than most vendor tables: a fresh git worktree per task, the reply parsed back into files, and pass defined as the verifier exiting 0 on the first attempt with no human edit and no retry [3]. Twenty for twenty under that rule [1] is the finding I would carry forward, while the ordering by seconds is the part I would re-measure.
The transfer condition sits in the harness cost. The post reports about five minutes of wall time per provider on the network [4]. Five OpenRouter calls at a 2.9-second mean is roughly 14.5 seconds, under 5% of that [6]; the write-up does not break down the remainder, though the harness creates a worktree and runs a test suite for every task [3]. A 24-second inference difference is decisive when a human waits on a single-shot edit and invisible when the agent loop already spends minutes in the verifier. These numbers describe single-turn prompts with whole files pasted in and one code block back [3]. If that is not your call shape, re-derive before you route traffic on it.
One caution about the ledger. Seven free tiers were attempted; two required a credit card, two returned 404s on their model IDs, and one turned out to be a paid aggregator [2]. That subtraction leaves two, and three genuinely free providers are reported [7]. The write-up also credits "all 4 free providers" when its own setup is three free plus one paid Claude control [8]. Nothing in the timings depends on either slip, but they set the precision you should read into a single-trial table.
Ranked by verification strength, evidence, and original report placement.
On the cli-flag task, NVIDIA took 61.7 seconds on a long thinking block for a multi-file change while OpenRouter finished in 2.7 seconds, which the author reports as a 23x speed difference on the same prompt and task.
A dev.to author tested 3 free LLM API providers plus one paid Claude plan on the same 5 real coding tasks; all 20 runs passed on the first attempt, with no provider failing any task.
The author tried to test 7 free-tier providers: two required a credit card on file (Cerebras and the Z.ai free tier), two had model IDs that returned 404s on the day of testing, and one (ClinePass) turned out to be a paid aggregator rather than a free tier.
Method: each task ran in a fresh git worktree branched from main; current file contents were read into the prompt with an instruction to output each modified file in a file:path code block; the response was parsed, new files written, and the task verifier run (existing test suite plus a content check). Pass means the verifier exits 0 on the first attempt, with no human edit and no retry.
The run covered 20 (provider, task) pairs with 1 trial each, at about 5 minutes of wall time per provider on the network.
OpenRouter's free tier, routing to the smaller open-weight MiniMax M3 model, had a 2.9-second mean wall time, which the author calls fast enough for interactive coding-agent work.
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · September 2, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Legible method, one sample per cell
Credit where it is due: dev.to shows the worktree isolation, the output contract, the pass rule, and a repo you can clone and re-run — more transparency than most free-tier roundups bother with. What the run cannot carry is the comparison it is asked to make. One trial per pair, five short tasks, and a perfect 20-for-20 score mean the harness never discriminated between providers on anything but the clock. And the post's arithmetic slips where it can be audited: seven candidates minus five stated disqualifications leaves two free tiers, yet three are tested and 'all 4 free providers' are congratulated.
One desk using it, providers moving underneath
Usage evidence stops at the author: he ran the suite, liked the Cline API enough to adopt it for future team benchmarks, and published the keys-and-clone instructions. No other team, product or deployment appears. The firmer adoption signal is on the supply side and it points the other way — the free tier everyone cited a year ago now asks for a card, two model IDs stopped answering, and OpenRouter's free pool swaps models without notice.
Headline outruns the sample
'Speed varies 10x' is doing more work than the data can support. The 9.3x mean spread comes from single calls, and the 23x showpiece is one 61.7-second thinking block against one 2.7-second run; remove that task from both providers and the gap shrinks to roughly 6x. The bigger stretch is the conclusion that the capability gap is now 'style, not correctness' — five docstring-and-flag tasks scored by an exit code cannot see correctness, and the author says himself that tool use went unmeasured. To his credit, he flags the personal-versus-team ceiling and lists what he did not test, which keeps this from being pure promotion.
Own repo, own dark horse
The post ends in a git clone of the author's own experiments repo, and its most enthusiastic recommendation is the one product that charges money — the Cline API, which he discovered during the test, calls a dark horse, and says he will keep using. No sponsorship or relationship is disclosed anywhere, and the raw results are shipped alongside the conclusions, so the pull here is reputational and habitual rather than commercial. Still, the cheaper-than-Anthropic-direct line arrives with no pricing worked out.
Whole method visible, nothing corroborated
We can read the entire experiment — provider list, prompts contract, pass rule, per-task timings — which makes this easy to assess and impossible to verify. Our reading of what the run shows is firm; the run's own claim on reality is the weak link, and the two internal contradictions we can check both resolved against the post. A second person cloning the repo would move this number more than any further reading of the text.