Build1 publisher3 min readPublished
Dual 3090s, no NVLink: the serving stack broke long before the model did
A week-long failure log on two RTX 3090s under WSL2 lands on one config at 170-210 tok/s. Everything before it died in dependency resolution, not in the math.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- The author spent a full week getting Qwen3.8-27B (hybrid GDN architecture, 48 linear-attention plus 16 full-attention layers, built-in MTP head) running on a dual RTX 3090 (24GB x 2) box with no NVLink, PCIe Gen4, Windows 10 plus WSL2 (Ubuntu-24.04).
- The configuration the author finally arrived at delivered 170-210 tok/s on code and JSON workloads.
- vLLM 0.25.1 (Docker) loads, but MTP speculation gives zero benefit on two GPUs without NVLink (56 vs 58 tok/s) and pushes first-token latency from 8.7s to 15s.
- vLLM 0.27.1 hits a missing libnvrtc.so.13 because nvidia-cuda-nvrtc installed as a 0.0.0a0 placeholder package; after force-reinstalling, flashinfer's sampling JIT demands the CUDA-13-only --host-stub-linkage-explicit while local nvcc is 12.8, and engine init crashes every time.
- The author's verdict: skip vLLM 0.27+ for Qwen3.8 on 3090s; the old Docker 0.25.1 works but gives no speedup.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
A developer published a full week's failure log from putting Qwen3.8-27B onto a dual RTX 3090 box with no NVLink, PCIe Gen4, on Windows 10 plus WSL2 (Ubuntu 24.04) [1]. The interesting part is not the 170-210 tok/s the final config reached on code and JSON output [2], it is that almost every dead end was in the serving stack, the checkpoint, or the CUDA toolchain, and not in the model or the GPUs.
vLLM went first. Version 0.25.1 in Docker loaded and ran, but the model's built-in multi-token prediction head bought nothing across two GPUs without NVLink: 56 tok/s with speculation against 58 without, and first-token latency rising from 8.7s to 15s [3]. That is a 2 tok/s regression paid for with 6.3 extra seconds of TTFT [1]. Version 0.27.1 never initialised at all: nvidia-cuda-nvrtc had installed as a 0.0.0a0 placeholder so libnvrtc.so.13 was missing, and after a force reinstall flashinfer's sampling JIT demanded a CUDA-13-only nvcc flag against a local 12.8 toolchain [4]. The author's verdict is to skip vLLM 0.27+ for this model on 3090s [5].
SGLang was worse, hanging during weight loading with the port never listening [6]. Dozens of parameter combinations changed nothing; comparing keys in model.safetensors.index.json showed the checkpoint was half-quantized, with the MTP head and some linear-attention layers stored as plain `weight` rather than `weight_packed`, which is why the equivalent Qwen3.6 checkpoint loads and this one dies [7]. Three flags plus a source patch were needed to get past the no-NVLink and WSL failure modes at all [8][9], and the result ran at roughly 10 tok/s [10].
The throughput came from abandoning the quantized MTP head. Because the INT4 quantisation had also quantised a head that ships in BF16, acceptance sat at 1.07-1.27, which the author frames as a quantisation problem rather than a tuning one [12]. An independent 1.4B draft model proposing 7-token blocks reached acceptance of 3.7-4.2, around four tokens per verification pass [11].
Two hard limits are worth writing down. GDN kernels need CUDA 13 to compile, where 12.8 hangs on load and 13.3 fails to build, leaving 13.0 as the only working version [13]. And flashinfer's GDN kernels require SM90+, so on SM86 silicon only the Triton linear-attention backend is available [14]; `--enable-torch-compile` also crashes on GDN [15].
Two officially recommended settings cost real throughput here. The `extra_buffer` mamba radix cache strategy took decode from 62 to 30 tok/s, roughly 52 percent, as a permanent charge against a rare failure [16][2]. A 2048 chunked-prefill size, sensible for concurrency, quadrupled prefill iterations on single-request 50-90K contexts [17]. The most banal finding is the most reusable: a dropped trailing backslash made bash truncate the launch command and silently discard every later argument, booting the server on defaults [18].
These are self-reported numbers from one machine, translated with LLM help, and unverified [22]. Watch whether vLLM 0.27+ stops requiring a matching local nvcc, and whether quantisers start shipping W4A16 checkpoints that leave the MTP head in BF16, which is the difference between speculation working and needing a second model in VRAM.