Nvidia's DGX Spark 64GB ships October 23 from Acer, ASUS, Dell, Gigabyte, HP and MSI, with the same GB10 chip as the 128GB model and half its memory. Teams comparing it with cloud token bills now start by asking how much of their workload fits in 64GB.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+40
- Incentives75
- Confidence40
NVIDIA says multi-agent systems burn up to 15 times the tokens of a standard chat, and Nemotron 3 Super is its open-weight attempt to make each of those tokens cheaper to produce. The efficiency figures come with NVIDIA's own hardware and its own predecessor as the baselines.
Reality
- Evidence38
- Adoption18
- Hype gap+34
- Incentives88
- Confidence57
A dev.to walkthrough of a 600GB NVFP4 model on discounted 8xH100 spot nodes traces the two crashes that arrive before the first prompt to a pip resolver replacing numpy and a KV cache sized for a million tokens.
Reality
- Evidence26
- Adoption14
- Hype gap+34
- Incentives
- Insufficient
- Confidence28
NVIDIA's submission runs the same Qwen3.6-27B as the llama.cpp reference on the same Jetson board and finishes 6.4x sooner. Most of the gap comes from prompt tokens the runtime never has to prefill.
Reality
- Evidence58
- Adoption18
- Hype gap+20
- Incentives85
- Confidence62
MLPerf Inference v6.1 preview submissions put Vera Rubin NVL72 at up to 3.7x GB300 on Qwen3-VL and up to 2.5x on DeepSeek-R1, on two different inference frameworks. The four-rack 99% scaling result is an offline number.
Reality
- Evidence45
- Adoption30
- Hype gap+35
- Incentives85
- Confidence55
The model averages 53 steps a run against SWE-1.7's 127, and Cognition says its mean rollout cost is 64% below Fable 5.1's on a leaderboard Cognition built, runs and grades. Reproducing that takes Devin's harness.
Reality
- Evidence30
- Adoption25
- Hype gap+30
- Incentives80
- Confidence45
Alibaba's largest open-weight release fits on a single eight-GPU node only because a community four-bit build takes it down to 1.2 TB. AWS publishes the vLLM config for it and skips the price.
Reality
- Evidence52
- Adoption22
- Hype gap+32
- Incentives78
- Confidence48
Nemotron 3.5 Lightning touches a tenth of its weights per token, which is what makes an agent loop plausible next to the sensors, but memory still has to hold all thirty billion and NVIDIA's own loop still escalates off-device.
Reality
- Evidence42
- Adoption
- Insufficient
- Hype gap+38
- Incentives84
- Confidence54
The zettaFLOPS figure is an estimate. The measured one, sitting in a GB300 example, is a 25 percent density gain, and it is paid for out of the safety margin.
Reality
- Evidence42
- Adoption15
- Hype gap+34
- Incentives78
- Confidence55
A capture of 45 gradient tensors from a 3M-parameter transformer overturned the author's own conditional claim. The Hadamard rotation the claim rested on moves the numbers by under 4%.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+12
- Incentives35
- Confidence38
A published NVFP4 and speculative-decoding config turns a 27B open-weights model into something you can try to serve. The 206.1 tokens per second figure is single-stream and unreplicated.
Reality
- Evidence52
- Adoption28
- Hype gap+14
- Incentives70
- Confidence46
NVIDIA says Alibaba's largest open-weight model serves over 4K tokens/sec/GPU and 350 tokens/sec/user in FP8 on a GB300 NVL72. That figure is the self-hosting floor, not a benchmark.
Reality
- Evidence32
- Adoption42
- Hype gap+34
- Incentives88
- Confidence44
NVIDIA's Nemotron 3.5 Lightning is now deployable from SageMaker JumpStart and fits on one GPU. The argument underneath it is about billing, not benchmarks.
Reality
- Evidence38
- Adoption22
- Hype gap+32
- Incentives82
- Confidence45
A 66 GB checkpoint becomes 22 GB and, NVIDIA says, up to 4x faster throughput, but only after distillation teaches the quantized student to live with its own noise.
Reality
- Evidence38
- Adoption20
- Hype gap+32
- Incentives85
- Confidence46
An August 2026 paper argues low-bit quantization-aware training converges high because its reconstruction step ignores which weights matter. The fix is reported to cost 1.4% of step time.
Reality
- Evidence28
- Adoption
- Insufficient
- Hype gap+34
- Incentives
- Insufficient
- Confidence27