Skip to content

Build1 publisher2 min readPublished

Four of five Qwen 3.8 27B engines finish within about 10% at 64 streams in g factor's tests

g factor's Qwen 3.8 27B benchmark has Together AI fastest at one stream, at 189.61 tok/s, while four of five engines finish within about 10% at 64 streams. Choosing a provider from these numbers starts with knowing how many streams the deployment will run at once.

The Engineer · Build desk

Illustration accompanying Four of five Qwen 3.8 27B engines finish within about 10% at 64 streams in g factor's tests

What happened

  • g factor ran the tests over several weeks on its own gft-studio platform, and its own engine was measured alongside Together AI, Fireworks AI, Nebius and Doubleword.
  • The post credits the single-stream lead to tensor parallelism inside one node and reports that tensor parallelism stretched across nodes stalls.
  • Each request sent about 564 input tokens and asked for exactly 128 output tokens at temperature 0 and seed 42, streamed over public HTTPS with AIPerf 0.12.0.
  • The post reports that four-token multi-token prediction (MTP4) speculative decoding pushed dual-H100 throughput to 769 tok/s.
  • Nebius is named among the providers tested but has no row in the five-engine throughput chart of Together, Fireworks FP8, Doubleword, vanilla vLLM and g factor.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • contradiction g factor says a meaningful benchmark must eliminate confounding variables, yet its chart sets an FP8 Fireworks row and an MTP4 g factor row beside vanilla vLLM, so gaps between rows can reflect quantization and speculative decoding as well as the engine.
  • decision Teams serving mostly one or two interactive streams at a time have to ask each provider how the model is split across GPUs, because tensor parallelism paid off here only inside a single node.
  • cost Operators budgeting a reasoning workload still need their own test of how each gateway counts reasoning tokens, since these 128-token runs leave that cost unmeasured.

TP2 and DP2 use the same two GPUs in different ways. Under TP2, individual weight matrices are split across both GPUs over NVLink, so both GPUs work on every token [5]. Under DP2, two independent replicas sit behind a load balancer [6]. A single request on DP2 lands on one replica and decodes on one GPU, while the other GPU waits for a second request [1]. g factor wrote that hardware interconnects "dictate whether Tensor Parallelism flies or grinds to a halt" [20].

The write-up opens with the observation that every inference provider claims to be "the fastest engine on Earth" [19]. On g factor's data, the answer changes with the stream count. Together's output throughput at two streams was 350.57 tok/s [7]. Against its single-stream figure, that is about 92% of linear scaling, or roughly 175 tok/s per stream once a second request is in flight [2].

The harness itself is careful work. The model is pinned to a single tokenizer revision [9]. Each low-concurrency cell ran 16 warmup requests, then 60 measured requests, three times over, for 180 measured requests [9]. Cells at c16, c32 and c64 ran 64 warmups and 256 measured requests per repeat [10]. The hardware is held less tightly than the method. The chart caption describes it as "2x H100 (or the provider's equivalent)" [11]. The prefix-caching result of more than 1,140 tok/s comes from structured prompts [18].

Time to first token runs from HTTP request submission to the first stream chunk [16], over public HTTPS in this harness [14]. Every TTFT figure therefore includes the network path to each provider [5]. g factor wrote that calling an Oregon server from a London laptop tells you "about transit latency, not engine throughput" [15]. The harness section does not state where its own client ran [21].

g factor's introduction also warns that "subtle gateway interpretations of reasoning tokens can quietly balloon your generation budgets" [17]. Every measured request asked for exactly 128 output tokens [14]. No row in the throughput chart shows how a gateway handles a reasoning trace longer than that [3].

What to watch

  • g factor's B200-versus-H100 and MTP4-versus-MTP8 results, which the post says it covers, and whether Nebius gets a row in the throughput chart.
  • A rerun with reasoning-length outputs that measures how each provider's gateway counts and bills reasoning tokens.
  • An independent reproduction of Together's 189.61 tok/s single-stream figure from a client in a stated region.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories