Science1 publisher3 min readPublished
The real disclosure in Qwen3.8-Max is the rack: 2.4T open weights, 72 GPUs, 4K tokens/sec
NVIDIA says Alibaba's largest open-weight model serves over 4K tokens/sec/GPU and 350 tokens/sec/user in FP8 on a GB300 NVL72. That figure is the self-hosting floor, not a benchmark.
The Scientist · Science desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Alibaba released the open weights for Qwen3.8-2.4T-A95B (Qwen3.8-Max), its largest open-weight model, bringing near-frontier capabilities to the open ecosystem.
- The model has 2.4T total parameters with 95B activated per token, in a fine-grained mixture-of-experts architecture.
- The architecture is a fine-grained MoE with a hybrid of full and linear attention, a context window up to one million tokens, and output length up to 128K, designed for demanding reasoning and agentic workloads.
- Deploying a 2.4T parameter open-weight model requires data-center-scale accelerated compute, and inference at this scale depends on extreme co-design across chips, system architecture, and software.
- Without additional model tuning, the model achieves throughput of over 4K tokens per second per GPU and over 350 tokens per second per user on NVIDIA GB300 NVL72 in FP8 precision on Day 0.
Compiled by The ScientistSomething wrong?How this is made
Why it matters
Alibaba has published open weights for Qwen3.8-2.4T-A95B, branded Qwen3.8-Max, a fine-grained mixture-of-experts model with 2.4 trillion total parameters and 95 billion activated per token [1][2]. NVIDIA, in a Day-0 post, says it runs at more than 4,000 tokens per second per GPU and more than 350 tokens per second per user on a GB300 NVL72 in FP8 precision, without additional model tuning [5]. The second sentence is the operationally useful one: it prices the entry ticket for self-hosting frontier-class open weights.
The unit of deployment here is a rack, not a server. GB300 NVL72 puts 72 Blackwell Ultra GPUs into a single NVLink domain with 130 TB/s of all-to-all bandwidth, which NVIDIA says removes the bottleneck that appears when expert routing traffic has to cross conventional networking [7]. That matters because a fine-grained MoE moves tokens between many small experts every layer; the interconnect, not the FLOPs, is what usually breaks first. NVIDIA states plainly that a 2.4T-parameter model requires data-center-scale compute and extreme co-design across chips, system architecture and software [4].
Do the arithmetic on the disclosed numbers and the shape of the deployment becomes clear. At 4,000 tokens per second per GPU across 72 GPUs, one rack produces roughly 288,000 tokens per second in aggregate [8]. If the per-GPU and per-user figures come from the same configuration, 4,000 divided by 350 is about 11 concurrent streams per GPU, or roughly 820 per rack [9]. Separately, 2.4 trillion parameters at one byte each is about 2.4 TB of weight storage before any KV cache, which is the arithmetic behind the rack-scale requirement [11].
The architecture is what makes any of this tractable. Only about 4 percent of the parameters are active per token [10], and NVIDIA argues serving cost therefore tracks active parameters rather than the full 2.4T, giving frontier-scale capacity for a fraction of a comparable dense model [20]. The model alternates full-attention and linear-attention layers, replacing a growing KV cache with a bounded recurrent state in the linear layers, which keeps compute and memory bounded out to a one-million-token context with up to 128K output [13][3]. NVIDIA frames this against agentic workloads, where accumulated tool outputs, retrieved documents and reasoning traces make attention and KV cache the binding constraints [14]. Built-in low/high/xhigh reasoning controls let developers trade inference depth for throughput per request [12], which is a scheduling lever as much as a quality one.
Treat the throughput numbers with the caution any vendor-published figure deserves. The post reports them without stating concurrency, batch size, or input and output sequence lengths [19], so they are a ceiling from a tuned rack rather than a planning number for mixed traffic. NVIDIA also says NVFP4 and further optimisation should improve on this over time [6], which is an admission that FP8 is not the end state.
What to watch: whether the NVFP4 results arrive with methodology attached; whether the open recipes in SGLang, vLLM and NVIDIA Dynamo [15] reproduce anything close to 350 tokens per second per user outside a single NVLink domain; and how hosted endpoints on DeepInfra, DigitalOcean, Fireworks AI, Modal and OpenRouter [16] price against the cost of owning a rack. Weights are on Hugging Face and ModelScope [17], and NeMo AutoModel supports full SFT or LoRA on the Day-0 checkpoints [18] for teams that intend to fine-tune rather than rent.