Skip to content

Topic

Mixture of Experts Architectures

A neural network design that activates only a few specialized "expert" subnetworks per input, scaling parameter count without raising compute cost.

Current stories

build3 publishers

Reflection sells Beam to sovereign-AI buyers on a 3-to-4x inference compute claim

Reflection AI says Beam, its 501-billion-parameter open-weight model, needs three to four times less compute than comparable open models. Buyers have only the company's figures until the weights ship later in October.

Perspective Coverage

3 publishers
Builder
Builder 38%
Operator
Operator 27%
Investor
Investor 35%

Reality

Evidence38
Adoption8
Hype gap+40
Incentives72
Confidence60
build6 publishers

Aleph Alpha's open Kolibri model routes each token through 3.46B of its 78.1B parameters

Aleph Alpha released Kolibri, an Apache 2.0 German-English model that activates 3.46B of its 78.1B parameters per token. Each token costs about as much compute as a small model, yet a team hosting it in Europe still has to fit every expert in memory.

Perspective Coverage

6 publishers
Builder
Builder 47%
Operator
Operator 37%
Investor
Investor 16%

Reality

Evidence72
Adoption
Insufficient
Hype gap+15
Incentives65
Confidence70
build1 publisher

Moonshot publishes open weights for its trillion-parameter K2.7-Code model

Moonshot AI released open weights for Kimi K2.7-Code, a trillion-parameter coding model that activates 32 billion parameters per token. Its headline gains come from Moonshot's own benchmarks, so teams paying for proprietary agents have to measure it on their own code.

Publishers:dev.to

Reality

Evidence35
Adoption
Insufficient
Hype gap+20
Incentives55
Confidence40
build1 publisher

Patterned PEFT LoRA adapters run at the wrong scale in vLLM 0.30.0

vLLM 0.30.0 ignores the per-module rank and alpha patterns in PEFT LoRA adapters, and one test put the error against PEFT at 70 times the unpatterned baseline. The adapter loads without complaint, so only a config check or a side-by-side against PEFT will catch it.

Publishers:dev.to

Reality

Evidence64
Adoption
Insufficient
Hype gap+12
Incentives20
Confidence66
build2 publishers

Ai2's Olmo-core 3 gets 2.7x its old MoE throughput by keeping experts resident on GPUs

Ai2 released Olmo-core 3, an open mixture-of-experts training stack it has benchmarked at over one trillion total parameters. Its speedups are measured against Ai2's own earlier FSDP code, so teams on Megatron-Core must run their own comparison.

Perspective Coverage

3 publishers
Builder
Builder 55%
Operator
Operator 25%
Investor
Investor 20%

Reality

Evidence50
Adoption
Insufficient
Hype gap+20
Incentives55
Confidence60
build3 publishers

DeepSeek's new encoder-decoder splits inference into an 8B prefill and a 16B decode

V4.1-Flash retires the V4 Pro line and carries two active-parameter counts, 763B total with 8B on input tokens and 16B on output, so one sizing number no longer covers both phases of a request. Baseten had it running on day zero.

Publishers:businesstimes.com.sgdev.tolatent.space

Perspective Coverage

3 publishers
Builder
Builder 40%
Operator
Operator 28%
Investor
Investor 32%

Reality

Evidence60
Adoption35
Hype gap+25
Incentives40
Confidence58

Earlier coverage

  1. A CreateAIBenchmarkJob call ramps a SageMaker endpoint from 64 to 1,024 concurrent requests

    Build · September 22, 2026 · 1 publisher

  2. White Circle publishes the harness behind Halo's 2.3x TRL throughput claim

    Build · September 21, 2026 · 1 publisher

  3. Filling GLM-5.3-Flash's million-token window costs three times its per-task benchmark price

    Build · September 20, 2026 · 2 publishers

  4. Three US providers host Moonshot's Kimi K3 at a tenth the cost of going to the source

    Invest · September 20, 2026 · 1 publisher

  5. NVIDIA's Nemotron 3 Super routes each token through a tenth of its 120 billion parameters

    Science · September 18, 2026 · 1 publisher

  6. Vals put Hy4 Preview first among open-weight models on code migration at $3.41 a test

    Build · September 17, 2026 · 1 publisher

  7. Atria Dawn's own team rated a third of its finished AI-assisted tasks infeasible without the agent

    Build · September 17, 2026 · 1 publisher

  8. Fireworks' own DeepSWE numbers put four coding models inside the noise band

    Product · September 17, 2026 · 1 publisher

  9. Qwen3.5-9B's 262K context window would consume the whole 8GB budget in KV cache

    Build · September 17, 2026 · 1 publisher

  10. An hour-long rollout and a minutes-long training step pushed Periodic Labs onto two GPU pools

    Build · September 16, 2026 · 1 publisher

  11. A 30B model that activates 3B still has to keep all 30B in VRAM

    Build · September 15, 2026 · 1 publisher

  12. A memory-mapped n-gram table still misses a quarter of page reads at 16 MB resident

    Build · September 14, 2026 · 1 publisher

  13. Top-k gating decouples DeepSeek-V3's 671B parameter count from its 37B of per-token compute

    Build · September 14, 2026 · 1 publisher

  14. NVIDIA's unoptimized DeepSeek-V3 baseline spends 84% of kernel time moving tokens between GPUs

    Build · September 14, 2026 · 1 publisher

  15. Hetzner's free inference experiment caps output at 60,000 tokens a minute

    Build · September 13, 2026 · 1 publisher

  16. DeepSeek's V4.1-Flash reads a million-token prompt on 8B active parameters

    Leadership · September 12, 2026 · 1 publisher

  17. Edge0's 35B would need 22 GB/s of SSD bandwidth without the file cache

    Build · September 11, 2026 · 2 publishers

  18. Bartowski broke tensors one at a time to find where GGUF bits belong

    Build · September 11, 2026 · 1 publisher

  19. DeepSeek reroutes V4-Pro API traffic to a smaller model on September 14

    Product · September 11, 2026 · 1 publisher

  20. Two open-weight launches clear frontier-tier numbers on vendor engineering claims

    Build · September 10, 2026 · 1 publisher

  21. NVFP4 squeezes Qwen3.8's 2.4 trillion weights onto eight B300s at 150 GB a GPU

    Build · September 9, 2026 · 1 publisher

  22. DeepSeek's V4 preview cuts million-token KV cache to a tenth of V3.2's

    Leadership · September 8, 2026 · 1 publisher

  23. Nex-AGI's runnable N2.5 tiers ask for two H100s or sixteen H200s

    Build · September 8, 2026 · 1 publisher

  24. Twelve of 500 byte-identical temperature-0 requests came back different

    Build · September 7, 2026 · 1 publisher

  25. Alibaba ships the Qwen4 architecture as open weights before the flagship exists

    Build · August 28, 2026 · 5 publishers

  26. GLM-5.3-Flash benchmarks its tenth-of-the-price claim against its own predecessor

    Leadership · September 5, 2026 · 1 publisher

  27. Ant Ling's Ling-3.0-flash-VL adds 1M-token context and a separate 32-frame video cap

    Build · September 4, 2026 · 1 publisher

  28. DeepSeek V4 moves the coding-model decision into the finance column

    Build · September 1, 2026 · 1 publisher

  29. Thirty-nine retries fit inside the price gap between GLM-5.3-Flash and Opus 4.8

    Build · August 31, 2026 · 1 publisher

  30. The sparse-model bill arrives at serving time, and it is paid in collectives

    Build · August 21, 2026 · 1 publisher