Reflection AI says Beam, its 501-billion-parameter open-weight model, needs three to four times less compute than comparable open models. Buyers have only the company's figures until the weights ship later in October.
Perspective Coverage
3 publishers
- Builder
- Builder 38%
- Operator
- Operator 27%
- Investor
- Investor 35%
Reality
- Evidence38
- Adoption8
- Hype gap+40
- Incentives72
- Confidence60
Reflection AI's 501-billion-parameter Beam beats GLM 5.2 on SWE-bench Pro and trails it on Terminal-Bench, by the company's own scores. Its claimed compute saving covers only part of a serving bill, so planning around it as a US-built open model has to wait for outside tests.
Reality
- Evidence35
- Adoption8
- Hype gap+30
- Incentives75
- Confidence40
Aleph Alpha released Kolibri, an Apache 2.0 German-English model that activates 3.46B of its 78.1B parameters per token. Each token costs about as much compute as a small model, yet a team hosting it in Europe still has to fit every expert in memory.
Perspective Coverage
6 publishers
- Builder
- Builder 47%
- Operator
- Operator 37%
- Investor
- Investor 16%
Reality
- Evidence72
- Adoption
- Insufficient
- Hype gap+15
- Incentives65
- Confidence70
Strata runs a 125B-class Qwen model at about 100 tok/s on one RTX 4090 because its sparse MoE activates only about 6B parameters per token. The design moves the hardware bill to system RAM, and the 100 tok/s figure holds mainly for code and structured output.
Reality
- Evidence45
- Adoption15
- Hype gap+40
- Incentives
- Insufficient
- Confidence40
Moonshot AI released open weights for Kimi K2.7-Code, a trillion-parameter coding model that activates 32 billion parameters per token. Its headline gains come from Moonshot's own benchmarks, so teams paying for proprietary agents have to measure it on their own code.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+20
- Incentives55
- Confidence40
Salvatore Sanfilippo's open-source ds4 engine runs a short list of large open-weight models locally by compressing their routed experts to about two bits. Even compressed, the supported builds need high-memory Macs or GPU systems that most people do not already own.
Reality
- Evidence55
- Adoption30
- Hype gap+20
- Incentives
- Insufficient
- Confidence60
vLLM 0.30.0 ignores the per-module rank and alpha patterns in PEFT LoRA adapters, and one test put the error against PEFT at 70 times the unpatterned baseline. The adapter loads without complaint, so only a config check or a side-by-side against PEFT will catch it.
Reality
- Evidence64
- Adoption
- Insufficient
- Hype gap+12
- Incentives20
- Confidence66
Ai2 released Olmo-core 3, an open mixture-of-experts training stack it has benchmarked at over one trillion total parameters. Its speedups are measured against Ai2's own earlier FSDP code, so teams on Megatron-Core must run their own comparison.
Perspective Coverage
3 publishers
- Builder
- Builder 55%
- Operator
- Operator 25%
- Investor
- Investor 20%
Reality
- Evidence50
- Adoption
- Insufficient
- Hype gap+20
- Incentives55
- Confidence60
VIDRAFT says POCKET-Darwin-180B, a 111 GB 4-bit GGUF build of its 180B mixture-of-experts model, runs in llama.cpp on about $1,400 of consumer hardware. The accuracy evidence so far is one MMLU-Pro comparison that VIDRAFT reports itself.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+45
- Incentives60
- Confidence35
Xiaomi's public dashboard put the MiMo 2.6 Pro reinforcement-learning run at $1.05 million after about 51 hours, roughly $20,500 an hour. Its restart notes and token count give other teams an all-in reference for pricing their own RL runs.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+15
- Incentives50
- Confidence55
Fudan University researchers studied 700-plus task logs from the Atria Dawn build and found humans made 85.5 percent of method and parameter calls. For teams copying the setup, output is capped by how fast people can decide and review what agents produce.
Reality
- Evidence50
- Adoption30
- Hype gap+10
- Incentives40
- Confidence45
Developer xbill9 rebuilt Google's QAT Gemma 4 26B as int4 and fit it on one TPU v6e with 53,888 tokens of KV cache, against 3,456 for RedHat's FP8 build. Throughput nearly doubles, with accuracy checked on one classification suite.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+15
- Incentives40
- Confidence55
Moonshot published 2.8 trillion open weights. At four bits per parameter that is about 1.4TB resident before any cache, which rules out the eight-way H100 node most teams assume.
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+10
- Incentives40
- Confidence58
Hy4 preview carries 2.6 times Hy3's parameters and 3.9 times its context window. Tencent is giving the weights away. That means the build column in next quarter's model budget gets priced in accelerator memory rather than in tokens.
Publishers:simonwillison.net · tencent.com Reality
- Evidence62
- Adoption18
- Hype gap+30
- Incentives65
- Confidence60
AWS says running DeepEP over its EFA network on EKS gives mixture-of-experts reinforcement learning 40% more throughput. Whether that reaches another cluster depends on how much of each training step goes to expert traffic between nodes.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+30
- Incentives80
- Confidence35
V4.1-Flash retires the V4 Pro line and carries two active-parameter counts, 763B total with 8B on input tokens and 16B on output, so one sizing number no longer covers both phases of a request. Baseten had it running on day zero.
Publishers:businesstimes.com.sg · dev.to · latent.space Perspective Coverage
3 publishers
- Builder
- Builder 40%
- Operator
- Operator 28%
- Investor
- Investor 32%
Reality
- Evidence60
- Adoption35
- Hype gap+25
- Incentives40
- Confidence58
Moonshot AI's Kimi Linear reports 75% less KV cache and 6.3x faster decoding than MLA at 1M tokens, using three linear layers per attention layer. The speedup falls to parity at 4k tokens, so the saving goes to traffic that runs at hundreds of thousands of tokens.
Reality
- Evidence40
- Adoption35
- Hype gap+25
- Incentives60
- Confidence45
Two agentic post-training jobs, a 1.02-trillion-parameter model and a 309-billion cousin, ran with step time, reward and infrastructure error rates on the open web. Both stopped near step 30, and both models are still unreleased.
Publishers:rajeshparikh.substack.com
Reality
- Evidence34
- Adoption14
- Hype gap+31
- Incentives62
- Confidence41
Flash's off-peak input price is under a quarter of what V4-Pro cost, and on DeepSeek's own table it beats the old Pro checkpoint on Terminal-Bench, but it scores 36.8 on Humanity's Last Exam and no V4.1-Pro has a date.
Reality
- Evidence42
- Adoption52
- Hype gap+30
- Incentives72
- Confidence40
Artificial Analysis scores the new Xiaomi model first among open weights and twelve points behind Claude Opus 5.5, and the cheaper Flash tier is the one an operator should put in front of a real queue.
Reality
- Evidence55
- Adoption30
- Hype gap+15
- Incentives72
- Confidence55
Earlier coverage
- A CreateAIBenchmarkJob call ramps a SageMaker endpoint from 64 to 1,024 concurrent requests
Build · September 22, 2026 · 1 publisher
- White Circle publishes the harness behind Halo's 2.3x TRL throughput claim
Build · September 21, 2026 · 1 publisher
- Filling GLM-5.3-Flash's million-token window costs three times its per-task benchmark price
Build · September 20, 2026 · 2 publishers
- Three US providers host Moonshot's Kimi K3 at a tenth the cost of going to the source
Invest · September 20, 2026 · 1 publisher
- NVIDIA's Nemotron 3 Super routes each token through a tenth of its 120 billion parameters
Science · September 18, 2026 · 1 publisher
- Vals put Hy4 Preview first among open-weight models on code migration at $3.41 a test
Build · September 17, 2026 · 1 publisher
- Atria Dawn's own team rated a third of its finished AI-assisted tasks infeasible without the agent
Build · September 17, 2026 · 1 publisher
- Fireworks' own DeepSWE numbers put four coding models inside the noise band
Product · September 17, 2026 · 1 publisher
- Qwen3.5-9B's 262K context window would consume the whole 8GB budget in KV cache
Build · September 17, 2026 · 1 publisher
- An hour-long rollout and a minutes-long training step pushed Periodic Labs onto two GPU pools
Build · September 16, 2026 · 1 publisher
- A 30B model that activates 3B still has to keep all 30B in VRAM
Build · September 15, 2026 · 1 publisher
- A memory-mapped n-gram table still misses a quarter of page reads at 16 MB resident
Build · September 14, 2026 · 1 publisher
- Top-k gating decouples DeepSeek-V3's 671B parameter count from its 37B of per-token compute
Build · September 14, 2026 · 1 publisher
- NVIDIA's unoptimized DeepSeek-V3 baseline spends 84% of kernel time moving tokens between GPUs
Build · September 14, 2026 · 1 publisher
- Hetzner's free inference experiment caps output at 60,000 tokens a minute
Build · September 13, 2026 · 1 publisher
- DeepSeek's V4.1-Flash reads a million-token prompt on 8B active parameters
Leadership · September 12, 2026 · 1 publisher
- Edge0's 35B would need 22 GB/s of SSD bandwidth without the file cache
Build · September 11, 2026 · 2 publishers
- Bartowski broke tensors one at a time to find where GGUF bits belong
Build · September 11, 2026 · 1 publisher
- DeepSeek reroutes V4-Pro API traffic to a smaller model on September 14
Product · September 11, 2026 · 1 publisher
- Two open-weight launches clear frontier-tier numbers on vendor engineering claims
Build · September 10, 2026 · 1 publisher
- NVFP4 squeezes Qwen3.8's 2.4 trillion weights onto eight B300s at 150 GB a GPU
Build · September 9, 2026 · 1 publisher
- DeepSeek's V4 preview cuts million-token KV cache to a tenth of V3.2's
Leadership · September 8, 2026 · 1 publisher
- Nex-AGI's runnable N2.5 tiers ask for two H100s or sixteen H200s
Build · September 8, 2026 · 1 publisher
- Twelve of 500 byte-identical temperature-0 requests came back different
Build · September 7, 2026 · 1 publisher
- Alibaba ships the Qwen4 architecture as open weights before the flagship exists
Build · August 28, 2026 · 5 publishers
- GLM-5.3-Flash benchmarks its tenth-of-the-price claim against its own predecessor
Leadership · September 5, 2026 · 1 publisher
- Ant Ling's Ling-3.0-flash-VL adds 1M-token context and a separate 32-frame video cap
Build · September 4, 2026 · 1 publisher
- DeepSeek V4 moves the coding-model decision into the finance column
Build · September 1, 2026 · 1 publisher
- Thirty-nine retries fit inside the price gap between GLM-5.3-Flash and Opus 4.8
Build · August 31, 2026 · 1 publisher
- The sparse-model bill arrives at serving time, and it is paid in collectives
Build · August 21, 2026 · 1 publisher