build1 distinct publisher Moonshot published 2.8 trillion open weights. At four bits per parameter that is about 1.4TB resident before any cache, which rules out the eight-way H100 node most teams assume.
Publishers:dev.to
Reality
- Evidence42
- Adoption20
- Hype gap+8
- Incentives
- Insufficient
- Confidence38
NVIDIA says Alibaba's largest open-weight model serves over 4K tokens/sec/GPU and 350 tokens/sec/user in FP8 on a GB300 NVL72. That figure is the self-hosting floor, not a benchmark.
Publishers:developer.nvidia.com
Reality
- Evidence32
- Adoption42
build1 distinct publisher The update adds a path selector and a two-tap convolution rather than layers, recovering most of the accuracy that tripling the drafter bought at 15.2% latency, by the vendor's own numbers.
Publishers:runtimewire.com
Reality
- Evidence54
- Adoption66
build1 distinct publisher A dev.to explainer on local inference benchmarks makes a point worth pinning up: tokens per second is a function of how many users you tested with, not a property of the hardware.
Publishers:dev.to
Reality
- Evidence34
- Adoption
- Insufficient
- Hype gap
build1 distinct publisher A 9B MIT-licensed coding model that reportedly matches a 31B rival on SWE-Bench Verified is still unusable as a Claude Code backend, because the runtime never turns its tool-call XML into a file write.
Publishers:dev.to
Reality
- Evidence44
- Adoption18
build1 distinct publisher A Kubernetes serving guide makes a point most teams skip: the YAML you reviewed is identical across five engines, and everything it hides is what breaks in production.
Publishers:dev.to
Reality
- Evidence30
- Adoption28
build1 distinct publisher A week-long failure log on two RTX 3090s under WSL2 lands on one config at 170-210 tok/s. Everything before it died in dependency resolution, not in the math.
Publishers:dev.to
Reality
- Evidence38
- Adoption18
build1 distinct publisher A dev.to guide argues local LLM capacity planning collapses into one napkin equation. Run it first and the hardware shortlist writes itself, tier names and all.
Publishers:dev.to
Reality
- Evidence42
- Adoption20
build1 distinct publisher An rmcp walk-through concedes its tools are I/O-bound at 100-500ms per AWS call, then makes the rewrite argument on footprint, dependency isolation and schema drift instead.
Publishers:dev.to
Reality
- Evidence34
- Adoption18
build6 distinct publishers Muse Glimmer ships as Apache 2.0 weights sized for a 24GB card. Muse Spark 1.2 stays on Muse Code and the Meta Model API. Plan capacity for two tiers, not one.
Publishers:latent.space · mezha.net · runtimewire.com · simonwillison.net · theneuron.ai · thenewstack.io
Perspective Coverage
6 publishers
- Builder
- Builder 38%
- Operator
- Operator 31%
- Investor
- Investor 31%
build1 distinct publisher Alibaba scheduled a 27-billion-parameter vision-language model for August 14, alongside an already-published 2.4-trillion-parameter MoE. The smaller file is the consequential one.
Publishers:runtimewire.com
Reality
- Evidence30
- Adoption10
Forty-two Intel advisories cover privilege escalation in Xeon, TDX and CSME/SPS, plus medium-severity bugs across a wide set of AI and ML packages. AMD's five advisories add twelve more.
Publishers:securityweek.com
Reality
- Evidence58
- Adoption30