build1 publisher
Firing 4.8% of the weights per token still leaves 125GB to keep resident
Alibaba's Qwen3.8-Flash-Next preview activates 6B of its 125B parameters per token. Per-token compute drops to under a quarter of the dense 27B's, and about 125GB of weights still has to stay on device.
Publishers:dev.to
Reality
- Evidence34
- Adoption18
- Hype gap+15
- Incentives60
- Confidence42