Build1 distinct publisher3 min readUpdated
A single-box test in Japan put 76 tokens/s next to 4.4 tokens/s, then found the deciding variable elsewhere: which models held the output format and which invented reassurance.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
A Microsoft MVP based in Japan ran one identical Japanese prompt through seven local models on a single NVIDIA DGX Spark and published the timings next to the full, unedited answers [1]. The two rankings that came out of it, speed and shippability, are not the same list, and the second one is the one that determines whether the output can go in front of a customer [14][9].
The rig was a GB10 in a DGX Spark with roughly 121GiB of unified memory visible to the OS, running Ollama 0.32.9 [2]. Conditions were pinned across all seven: temperature 0, seed 42, 8192 context, a 320-token generation budget, thinking off [3]. Each model got one cold run after being unloaded, then three warm runs, with the warm average as the primary number [4]. The author deliberately did not tune each model toward its own recommended settings, on the grounds that doing so mixes configuration differences into model differences, and reserved separate diagnostics for any model that could not answer under the common conditions [5].
The task was narrow and checkable: a small or midsize business wants generative AI without sending customer data to an external cloud, list three suitable tasks, one line each in the form "Task name: reason / caveat," then a short adoption verdict, with product names and law names prohibited [6].
Warm averages: Qwen3.5 35B at 1.49s and 76.27 tokens/s, format held [7]; GLM-4.7-Flash at 1.94s and 64.64 tokens/s, format broken [8]; Qwen3.6 35B-A3B at 2.12s and 44.72 tokens/s, format held [9]; GPT-OSS 120B at 8.98s with an empty displayed answer [10]; Nemotron 3 Super 120B-A12B, deployed as the large quantized build, at 9.09s and 20.36 tokens/s, format held [1][11]; Gemma 4 31B at 9.81s and 10.41 tokens/s, format held [12]; Qwen2.5 72B at 21.68s and 4.42 tokens/s, format held [13].
That is a 14.6x spread in latency and a 17.3x spread in throughput [1][2], and none of it predicts the outcome that matters. Five of seven held the format, one broke it, one returned nothing [3]. The failures sit at the fast end: the second-fastest model broke format and the fourth-fastest spent 8.98s producing an empty field, while the slowest model in the set complied [9]. Parameter count is no better as a proxy. Gemma 4 31B was slower than the 120B-A12B Nemotron build [5], and the MoE Qwen3.6 35B-A3B ran at 44.72 tokens/s against 76.27 for dense Qwen3.5 35B [7].
The interesting failure is the one that passed. Qwen3.5 was fastest by a wide margin, 6.1x quicker than Nemotron [4], and the author's point is that its actual text is why speed alone is a dangerous basis for choosing [16]. Its first line offered customer support log analysis with "zero risk of confidential data leakage" as the reason [17]. That is an absolute safety claim, unprompted, in a document about handling customer data. It formats correctly and it is not defensible.
Worth watching: the author poses but, in this piece, frames as open the question of whether Nemotron 3 Super's 87GB of weights earn their keep against a 23GB-class model, a 3.8x difference in footprint [18][6]. Also note the provenance. The article was drafted by Claude Code and Codex agents from the measurement data and then reviewed and edited by the author [19], and the measurement records and prior conditions are kept in Obsidian via a plugin the author maintains [20]. Single operator, single box, one prompt. The method is worth copying; the rankings are not worth quoting as general truth.
Follow any of these and your For You feed starts watching them — no settings page required.
Ranked by verification strength, evidence, and original report placement.
The author, a Microsoft MVP based in Japan, ran the exact same Japanese question through 7 major models on a single NVIDIA DGX Spark, including NVIDIA's Nemotron 3 Super 120B-A12B, for which the large quantized version was deployed, and published the raw measurements plus every full answer unedited.
The test used the GB10 in a DGX Spark; unified memory available from the OS is about 121GiB, and Ollama is version 0.32.9.
Fixed comparison conditions were temperature 0, seed 42, context 8192, max generation budget 320 tokens, and thinking OFF.
Each model got one cold run after unloading, followed by 3 warm runs; the primary comparison was the average response time across the 3 warm runs.
The author states that optimizing each model individually toward its own recommended settings would blend configuration differences with model differences, so the primary comparison stayed on common conditions and separate diagnostics were run only to check the cause for any model that could not answer under those conditions.
The prompt described a small or midsize business wanting to adopt generative AI without sending customer data to an external cloud AI, and asked for 3 tasks suited to local AI, each item on a single line in the form "Task name: reason / caveat," followed by a short adoption verdict; inventing product names or law names was prohibited.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Documented single-operator run, no replication
The measurement protocol is unusually explicit for a blog benchmark: fixed decoding settings, a stated seed, a cold-run-then-three-warm-runs procedure, named runtime version, memory ceiling, and full unedited model outputs. That makes the numbers checkable in principle. It is still one prompt, in one language, on one machine, by one author, with no run-to-run variance reported, no root cause given for the empty 120B output, and a body that breaks off before the Nemotron conclusion. Nothing in the cluster corroborates it independently.
One practitioner box, no third-party uptake
Adoption evidence is limited to a single practitioner's own environment: one DGX Spark, one Ollama version, seven publicly available models pulled locally, plus the author's self-built tooling. There is no organizational deployment, no user counts, no downstream party acting on the results, and no other publisher in the cluster reporting comparable runs.
Claims slightly under the measurements
The framing runs against the usual inflation. The author leads with a speed table that would support a punchy leaderboard claim, then argues the ranking does not decide the choice, flags his fastest model's 'zero risk' assurance as an overstatement, marks his preferred pick as editorial judgment on one question, and discloses AI drafting and his own tooling. Slightly negative rather than zero because the measured failure modes — an empty 120B answer and a broken output format — carry more operational weight than the cautious presentation gives them.
Disclosed self-promotion around own tooling
The author has a visible interest in the story being read: a Substack cross-post, an MVP practitioner brand, and two of his own products — the Ebi Workspace plugin and the open-source Ebi Agent Chat Relay — presented as the workflow behind the measurements. That is promotional but disclosed, and the incentive does not obviously bend the numbers, which include unflattering results for popular models. No vendor sponsorship, hardware loan or paid relationship with NVIDIA or any model provider is claimed or evidenced.
Internally consistent, externally unverified
Confidence is moderate: the arithmetic and the derived ratios follow directly from published figures, and the methodological disclosures are unusually complete for the format. It is held down by single-source, single-run-set provenance, a truncated body that omits the Nemotron verdict, unspecified quantization details, and the absence of any corroborating publisher in the cluster.
build
NVIDIA put a number on agent skills: 300+ verified, two harnesses, baselines under 50/1001 distinct publisher
build
A 12MB Go binary bets agent cost control is cache stickiness, not a dashboard1 distinct publisher
build
Block's Berd makes a duller argument than its mascots: show the agent's context as product state1 distinct publisher
build
Waku 0.1.0 bets the product is the control plane, not another coding agent1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 15, 2026