Moonshot AI released open weights for Kimi K2.7-Code, a trillion-parameter coding model that activates 32 billion parameters per token. Its headline gains come from Moonshot's own benchmarks, so teams paying for proprietary agents have to measure it on their own code.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+20
- Incentives55
- Confidence40
Moonshot published 2.8 trillion open weights. At four bits per parameter that is about 1.4TB resident before any cache, which rules out the eight-way H100 node most teams assume.
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+10
- Incentives40
- Confidence58
IBM's 3B, 8B and 30B dense models all get a thinking switch and native tool calling, but only the two larger ones get agentic RL, and the tuning mixture leans hard on software engineering.
Perspective Coverage
3 publishers
- Builder
- Builder 63%
- Operator
- Operator 28%
- Investor
- Investor 9%
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+25
- Incentives60
- Confidence72
Liquid AI shipped a 279.5-million-parameter draft model for LFM2.5-VL-3B on September 24, reporting 2.30x to 3.13x faster decoding on an Apple M5 Max. Image encoding and prompt prefill are unchanged, and the company's tests did not cover quantized deployments.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+15
- Incentives70
- Confidence60
NVIDIA's SWE-Serve scores the same 627 patches twice on 19 SGLang tasks, once with the live-serving tests and once without. The pass rate falls from 69.4% to 45.9%, and 242 of the 276 live tests came from SGLang itself.
Reality
- Evidence58
- Adoption30
- Hype gap−10
- Incentives70
- Confidence60
DigitalOcean's serving guide puts the cost of a long request on two mechanics, the prefill pass and the KV cache, and says the real choice is which mix of caching, prefix reuse and retrieval fits the budget.
Reality
- Evidence40
- Adoption
- Insufficient
- Hype gap+10
- Incentives80
- Confidence45
Halo adds expert and tensor parallelism to Hugging Face models and still saves SafeTensors that from_pretrained can load. Its best number, 9,009 tokens per second per GPU against TRL's 3,885, came from synthetic fixed-length sequences.
Reality
- Evidence45
- Adoption14
- Hype gap+22
- Incentives72
- Confidence56
AWS counted 13 SageMaker inference launches so far in 2026. The one that changes production behaviour most is a prioritized list of up to five instance types, each allowed its own model optimization settings.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+25
- Incentives85
- Confidence55
The designated successor to GenAI-Perf splits load generation from record processing so the benchmark client stops hitting Python's GIL. Teams holding a GenAI-Perf baseline inherit a port and a re-run.
Reality
- Evidence55
- Adoption20
- Hype gap+15
- Incentives75
- Confidence55
The number came out of one harness on OpenRouter at temperature 0.9 with a 64,000-token output cap, and it holds for a bounded translation task while the same model sits mid-table on open-ended terminal work.
Reality
- Evidence57
- Adoption28
- Hype gap+19
- Incentives62
- Confidence60
Periodic Labs has published the infrastructure account behind Neon, its trillion-parameter model, and the figures in it measure the training and serving harness down to a single 1.90-millisecond routing payload.
Reality
- Evidence58
- Adoption42
- Hype gap+12
- Incentives62
- Confidence58
A September 3 playbook traces speculative decoding's draft architectures from EAGLE-3 to DFlash, and grounds the case in a 70B model that decodes at 15 to 20 tokens a second on eight H100s.
Reality
- Evidence24
- Adoption31
- Hype gap+38
- Incentives38
- Confidence33
Ao Qu and collaborators open-sourced Reef on September 15th, an OpenAI-format inference server that stamps each response with a record ID, matches later feedback to that ID, and publishes a retrained artifact only if its evaluation stage accepts it.
Reality
- Evidence48
- Adoption10
- Hype gap+12
- Incentives32
- Confidence60
SemiAnalysis says its AgentX benchmark clocked Nvidia's Vera Rubin NVL72 at up to seven times Blackwell's token throughput per megawatt on pre-release software, and puts the profit gain at over two times per gigawatt.
Reality
- Evidence46
- Adoption29
- Hype gap+16
- Incentives67
- Confidence41
ROCm sits under PyTorch, vLLM and SGLang; Vulkan is what llama.cpp-class engines compile shaders against. The engine you already run narrows the choice to one, and the rest of the work is reading logs.
Reality
- Evidence52
- Adoption22
- Hype gap−8
- Incentives30
- Confidence48
A custom SGLang framework feeds teacher outputs into the loss while training runs. The reported eightfold speedup comes from a stack of separate optimizations, and the named multipliers reach about 5.5x on their own.
Reality
- Evidence45
- Adoption50
- Hype gap+15
- Incentives60
- Confidence50
Thinking Machines Lab argues that temperature-zero inference varies run to run not because GPUs are chaotic but because kernels change their reduction order with load, which makes bit-identical output an engineering target.
Publishers:thinkingmachines.ai
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap+12
- Incentives60
- Confidence55
Alibaba's 27B model fits a 32GB card with 15GB to spare. Tom's Hardware still had to pick between llama.cpp's full 262K window at half-hour prefill and a supported vLLM deployment capped at 32K.
Reality
- Evidence60
- Adoption38
- Hype gap+15
- Incentives55
- Confidence58
Qwen3.8-Flash-Next puts 36 Gated DeltaNet layers and 12 sparse-attention layers on Hugging Face, which means the retrieval budget Qwen4 will inherit is something you can measure against your own traces now.
Perspective Coverage
5 publishers
- Builder
- Builder 52%
- Operator
- Operator 28%
- Investor
- Investor 20%
Reality
- Evidence58
- Adoption52
- Hype gap+32
- Incentives76
- Confidence71
A cache hit answers 13 to 31 percent faster than a miss, and an ICML 2025 audit found seven of 17 live providers sharing that cache across users, though the one demonstrated 100 percent reconstruction ran against self-hosted vLLM and SGLang.
Reality
- Evidence38
- Adoption22
- Hype gap+34
- Incentives68
- Confidence46
Earlier coverage
- GLM-5.3-Flash benchmarks its tenth-of-the-price claim against its own predecessor
Leadership · September 5, 2026 · 1 publisher
- Peak-hour pricing pushes DeepSeek's new vision model past Gemini on the invoice test
Build · August 31, 2026 · 2 publishers
- DeepSeek's MIT-licensed V4-Pro hands API buyers a credible walk-away option
Leadership · August 30, 2026 · 1 publisher
- Four-bit weights leave 6 GB on a 24 GB card for KV cache and vision tensors
Build · August 29, 2026 · 1 publisher
- Intel puts its Arc GPU operating knowledge inside the coding agent already installed
Build · August 28, 2026 · 1 publisher
- Streaming tool-call deltas turn a base-URL swap into a per-model parser project
Build · August 28, 2026 · 1 publisher
- SPEED-Bench re-tests speculative decoding at the batch size you actually serve
Build · August 27, 2026 · 1 publisher
- Z.ai's cost-parity claim on Chinese accelerators rests on model design as much as silicon
Leadership · August 27, 2026 · 1 publisher
- Two buyers, one price: Patel says 2027's new compute is already half spoken for
Build · August 25, 2026 · 1 publisher
- Shadow engines cut LLM restart from 283 seconds to 7.3, and change what headroom is for
Build · August 25, 2026 · 1 publisher
- Nvidia's agent-workload lead scales with interactivity, and the second-source budget line does not
Leadership · August 25, 2026 · 1 publisher
- MiniMax maps H3 from one 24GB card to SGLang, and keeps the interpreter in-house
Build · August 24, 2026 · 1 publisher
- A benchmark that replays real agent sessions gives back less of the generational win
Build · August 24, 2026 · 1 publisher
- A 27B model reportedly beat a license check in 30 minutes. Nobody has seen the binary.
Build · August 23, 2026 · 1 publisher
- Ollama, vLLM, SGLang: the throughput ceiling is set by the queue, not the weights
Build · August 22, 2026 · 1 publisher
- SGLang's one-GPU Qwen3.8-27B recipe is the useful half of the release
Build · August 21, 2026 · 1 publisher
- The real disclosure in Qwen3.8-Max is the rack: 2.4T open weights, 72 GPUs, 4K tokens/sec
Science · August 20, 2026 · 1 publisher
- Inco AI's DFlash 2: 21% longer accepted drafts for 1.3% latency and 18.5M parameters
Build · August 19, 2026 · 1 publisher
- Your vLLM Manifest Would Boot SGLang Too, And That Is the Problem
Build · August 18, 2026 · 1 publisher
- Dual 3090s, no NVLink: the serving stack broke long before the model did
Build · August 18, 2026 · 1 publisher
- Qwen3.8's 27B dense checkpoint is the one operators can actually host
Build · August 14, 2026 · 1 publisher