Skip to content

Build1 publisher3 min readPublished

Seven local models, one prompt, one DGX Spark: the speed ranking decided nothing

A single-box test in Japan put 76 tokens/s next to 4.4 tokens/s, then found the deciding variable elsewhere: which models held the output format and which invented reassurance.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Seven local models, one prompt, one DGX Spark: the speed ranking decided nothing
Generated illustration

What happened

  • The author, a Microsoft MVP based in Japan, ran the exact same Japanese question through 7 major models on a single NVIDIA DGX Spark, including NVIDIA's Nemotron 3 Super 120B-A12B, for which the large quantized version was deployed, and published the raw measurements plus every full answer unedited.
  • The test used the GB10 in a DGX Spark; unified memory available from the OS is about 121GiB, and Ollama is version 0.32.9.
  • Fixed comparison conditions were temperature 0, seed 42, context 8192, max generation budget 320 tokens, and thinking OFF.
  • Each model got one cold run after unloading, followed by 3 warm runs; the primary comparison was the average response time across the 3 warm runs.
  • The author states that optimizing each model individually toward its own recommended settings would blend configuration differences with model differences, so the primary comparison stayed on common conditions and separate diagnostics were run only to check the cause for any model that could not answer under those conditions.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A Microsoft MVP based in Japan ran one identical Japanese prompt through seven local models on a single NVIDIA DGX Spark and published the timings next to the full, unedited answers [1]. The two rankings that came out of it, speed and shippability, are not the same list, and the second one is the one that determines whether the output can go in front of a customer [14][9].

The rig was a GB10 in a DGX Spark with roughly 121GiB of unified memory visible to the OS, running Ollama 0.32.9 [2]. Conditions were pinned across all seven: temperature 0, seed 42, 8192 context, a 320-token generation budget, thinking off [3]. Each model got one cold run after being unloaded, then three warm runs, with the warm average as the primary number [4]. The author deliberately did not tune each model toward its own recommended settings, on the grounds that doing so mixes configuration differences into model differences, and reserved separate diagnostics for any model that could not answer under the common conditions [5].

The task was narrow and checkable: a small or midsize business wants generative AI without sending customer data to an external cloud, list three suitable tasks, one line each in the form "Task name: reason / caveat," then a short adoption verdict, with product names and law names prohibited [6].

Warm averages: Qwen3.5 35B at 1.49s and 76.27 tokens/s, format held [7]; GLM-4.7-Flash at 1.94s and 64.64 tokens/s, format broken [8]; Qwen3.6 35B-A3B at 2.12s and 44.72 tokens/s, format held [9]; GPT-OSS 120B at 8.98s with an empty displayed answer [10]; Nemotron 3 Super 120B-A12B, deployed as the large quantized build, at 9.09s and 20.36 tokens/s, format held [1][11]; Gemma 4 31B at 9.81s and 10.41 tokens/s, format held [12]; Qwen2.5 72B at 21.68s and 4.42 tokens/s, format held [13].

That is a 14.6x spread in latency and a 17.3x spread in throughput [1][2], and none of it predicts the outcome that matters. Five of seven held the format, one broke it, one returned nothing [3]. The failures sit at the fast end: the second-fastest model broke format and the fourth-fastest spent 8.98s producing an empty field, while the slowest model in the set complied [9]. Parameter count is no better as a proxy. Gemma 4 31B was slower than the 120B-A12B Nemotron build [5], and the MoE Qwen3.6 35B-A3B ran at 44.72 tokens/s against 76.27 for dense Qwen3.5 35B [7].

The interesting failure is the one that passed. Qwen3.5 was fastest by a wide margin, 6.1x quicker than Nemotron [4], and the author's point is that its actual text is why speed alone is a dangerous basis for choosing [16]. Its first line offered customer support log analysis with "zero risk of confidential data leakage" as the reason [17]. That is an absolute safety claim, unprompted, in a document about handling customer data. It formats correctly and it is not defensible.

Worth watching: the author poses but, in this piece, frames as open the question of whether Nemotron 3 Super's 87GB of weights earn their keep against a 23GB-class model, a 3.8x difference in footprint [18][6]. Also note the provenance. The article was drafted by Claude Code and Codex agents from the measurement data and then reviewed and edited by the author [19], and the measurement records and prior conditions are kept in Obsidian via a plugin the author maintains [20]. Single operator, single box, one prompt. The method is worth copying; the rankings are not worth quoting as general truth.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories