build1 distinct publisher Moonshot published 2.8 trillion open weights. At four bits per parameter that is about 1.4TB resident before any cache, which rules out the eight-way H100 node most teams assume.
Publishers:dev.to
Reality
- Evidence42
- Adoption20
- Hype gap+8
- Incentives
- Insufficient
- Confidence38
NVIDIA says Alibaba's largest open-weight model serves over 4K tokens/sec/GPU and 350 tokens/sec/user in FP8 on a GB300 NVL72. That figure is the self-hosting floor, not a benchmark.
Publishers:developer.nvidia.com
Reality
- Evidence32
- Adoption42
build1 distinct publisher The update adds a path selector and a two-tap convolution rather than layers, recovering most of the accuracy that tripling the drafter bought at 15.2% latency, by the vendor's own numbers.
Publishers:runtimewire.com
Reality
- Evidence54
- Adoption66
build1 distinct publisher A Kubernetes serving guide makes a point most teams skip: the YAML you reviewed is identical across five engines, and everything it hides is what breaks in production.
Publishers:dev.to
Reality
- Evidence30
- Adoption28
build1 distinct publisher A week-long failure log on two RTX 3090s under WSL2 lands on one config at 170-210 tok/s. Everything before it died in dependency resolution, not in the math.
Publishers:dev.to
Reality
- Evidence38
- Adoption18
build1 distinct publisher Alibaba scheduled a 27-billion-parameter vision-language model for August 14, alongside an already-published 2.4-trillion-parameter MoE. The smaller file is the consequential one.
Publishers:runtimewire.com
Reality
- Evidence30
- Adoption10