Aleph Alpha released Kolibri, an Apache 2.0 German-English model that activates 3.46B of its 78.1B parameters per token. Each token costs about as much compute as a small model, yet a team hosting it in Europe still has to fit every expert in memory.
Perspective Coverage
4 publishers
- Builder
- Builder 47%
- Operator
- Operator 36%
- Investor
- Investor 17%
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap+15
- Incentives68
- Confidence66
Moonshot AI released open weights for Kimi K2.7-Code, a trillion-parameter coding model that activates 32 billion parameters per token. Its headline gains come from Moonshot's own benchmarks, so teams paying for proprietary agents have to measure it on their own code.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+20
- Incentives55
- Confidence40
Nebius bought Inferize, a 10-month-old Israeli startup with 17 staff, for an estimated $100 million to $150 million, according to Calcalist. The money goes on software that keeps inference GPUs from sitting idle, and Nebius is still adding data center capacity alongside it.
Perspective Coverage
3 publishers
- Builder
- Builder 38%
- Operator
- Operator 32%
- Investor
- Investor 30%
Reality
- Evidence55
- Adoption12
- Hype gap+30
- Incentives70
- Confidence60
vLLM 0.30.0 ignores the per-module rank and alpha patterns in PEFT LoRA adapters, and one test put the error against PEFT at 70 times the unpatterned baseline. The adapter loads without complaint, so only a config check or a side-by-side against PEFT will catch it.
Reality
- Evidence64
- Adoption
- Insufficient
- Hype gap+12
- Incentives20
- Confidence66
vLLM 0.30.0's weight cache fingerprints only safetensors headers, letting an engine on an altered Qwen3-0.6B run the daemon's original weights. Fine-tunes share their base model's headers, so a mix-up there would produce fluent wrong answers under a log line reporting success.
Reality
- Evidence64
- Adoption
- Insufficient
- Hype gap+8
- Incentives30
- Confidence62
Helion's autotuned GEMM beat vLLM's default CUTLASS and DeepGEMM backends on Hopper GPUs, by more than 10% throughput on some workloads, its authors report. The gain rests on per-shape tuning that can run for hours, a cost teams pay in place of kernel maintenance.
Reality
- Evidence35
- Adoption15
- Hype gap+10
- Incentives70
- Confidence40
Repacking Google's quantization-aware-trained Gemma 4 weights without re-rounding fits the 12B model on one TPU v5e chip at bf16-level accuracy. Adopting it means carrying patches to vLLM's TPU backend and working inside a 9,728-token KV cache.
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+15
- Incentives35
- Confidence55
KAIST and Seoul National University's AgSpec nearly doubles accepted draft length for coding agents by indexing files in the diff and JSON forms agents emit. It lives entirely in the retrieval index and leaves model weights untouched.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+15
- Incentives
- Insufficient
- Confidence40
g factor's Qwen 3.8 27B benchmark has Together AI fastest at one stream, at 189.61 tok/s, while four of five engines finish within about 10% at 64 streams. Choosing a provider from these numbers starts with knowing how many streams the deployment will run at once.
Reality
- Evidence48
- Adoption
- Insufficient
- Hype gap+18
- Incentives72
- Confidence45
SageMaker costs 1.40x a plain EC2 instance per hour to serve one Gemma 4 vLLM build on the same T4 or L4 GPU, a benchmark on dev.to finds. With decode speed matched within 2%, the extra 40% goes to the managed layer around the GPU.
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+5
- Incentives
- Insufficient
- Confidence60
Gemma 4 decodes on SageMaker's smallest GPU, a T4, at 0.8x an L4's speed with matching outputs from a patched vLLM image, a dev.to benchmark reports. The T4 is the cheaper choice per token only when it rents for under 80% of the L4's hourly rate.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+15
- Incentives35
- Confidence40
Google's Threat Intelligence Group counted 141 flaws exploited in the wild from January to August, while monthly disclosures doubled to 10,740. Patch teams do better sorting by that exploited set than by the total, though attackers now reach some public flaws within days.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+15
- Incentives60
- Confidence50
Repacking Gemma 4's QAT weights with 4-bit embedding tables made decode up to 1.39x faster on one SageMaker L4, according to a dev.to benchmark series. For teams serving Gemma 4 on vLLM, how the weights are stored becomes a setting to measure alongside model size.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+10
- Incentives
- Insufficient
- Confidence50
Llama 8B on an RTX 4060 Ti 16GB fell from 42.5 to 3.8 tokens a second with a fifth of the model in system RAM, according to a dev.to benchmark. For a local coding agent, that makes VRAM for weights plus context the first spec to check, ahead of bandwidth.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+5
- Incentives
- Insufficient
- Confidence40
Unsloth fixed Studio in version 2026.6.9 after Pillar Security showed that reading a malicious model's config.json could run an attacker's Python code. Unsloth disputed parts of the finding and declined to publish an advisory, so no CVE was assigned.
Reality
- Evidence48
- Adoption
- Insufficient
- Hype gap+10
- Incentives55
- Confidence52
Qwen Flash-Next NVFP4 ran in vLLM at 131,072 tokens of context and 16 sequences only after loader patches and cuts to its original targets. Its one load test covers only the bfloat16 KV-cache baseline, so operators on the later B12x stack have to run their own.
Reality
- Evidence45
- Adoption8
- Hype gap0
- Incentives
- Insufficient
- Confidence40
Developer xbill9 rebuilt Google's QAT Gemma 4 26B as int4 and fit it on one TPU v6e with 53,888 tokens of KV cache, against 3,456 for RedHat's FP8 build. Throughput nearly doubles, with accuracy checked on one classification suite.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+15
- Incentives40
- Confidence55
Moonshot published 2.8 trillion open weights. At four bits per parameter that is about 1.4TB resident before any cache, which rules out the eight-way H100 node most teams assume.
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+10
- Incentives40
- Confidence58
IBM's 3B, 8B and 30B dense models all get a thinking switch and native tool calling, but only the two larger ones get agentic RL, and the tuning mixture leans hard on software engineering.
Perspective Coverage
3 publishers
- Builder
- Builder 63%
- Operator
- Operator 28%
- Investor
- Investor 9%
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+25
- Incentives60
- Confidence72
Gemma 4 E2B's 4-bit QAT checkpoint decodes 2.05x faster than bf16 on one SageMaker L4, according to a dev.to benchmark. The swap also frees 18% of GPU memory for a 20% larger KV cache, though it runs only through the vLLM container because JumpStart lists no QAT build.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+8
- Incentives
- Insufficient
- Confidence55
Earlier coverage
- HyperPod holds the inference deployment until every target node has cached the weights
Build · September 10, 2026 · 1 publisher
- Kimi Linear's 75% KV cache cut comes from making three of every four layers recurrent
Build · September 25, 2026 · 1 publisher
- rocminfo reports one MI300X framebuffer three times, once per global memory pool
Build · September 16, 2026 · 1 publisher
- A 4-bit Gemma 4 26B on one L4 trails TypeSafe's Jev by 2.1 points overall
Build · September 23, 2026 · 1 publisher
- Folding 92 layers into one Pallas call put Kimi K3 at 709 tokens a second on TPU v7
Build · September 23, 2026 · 1 publisher
- Kaitchup traces Bonsai 2's 98.2% retention figure to unpacked weights on an H100
Build · September 22, 2026 · 1 publisher
- vLLM measured its portability layer at 3.4 percent below native throughput on an H100
Build · September 22, 2026 · 1 publisher
- A CreateAIBenchmarkJob call ramps a SageMaker endpoint from 64 to 1,024 concurrent requests
Build · September 22, 2026 · 1 publisher
- Crusoe's Series F values a $140 billion backlog at 22 cents on the dollar
Invest · September 22, 2026 · 1 publisher
- Shopify's continual learning loop cuts serving costs 96%; case study separately details judges calibrated from 25 annotated samples
Build · September 22, 2026 · 1 publisher
- VoxCPM2 trades the per-character speech bill for a GPU and a base-URL change
Build · September 21, 2026 · 1 publisher
- White Circle publishes the harness behind Halo's 2.3x TRL throughput claim
Build · September 21, 2026 · 1 publisher
- Paging the KV cache into 16-token blocks caps the waste at 960 KB per sequence
Build · September 21, 2026 · 1 publisher
- A resident Whisper model plus embedding model together burned 2 euros of GPU electricity across 30 days
Build · September 20, 2026 · 1 publisher
- Crusoe's $3.9bn round prices a contract book that runs five gigawatts ahead of delivery
Invest · September 19, 2026 · 2 publishers
- A local proxy convinces the ChatGPT desktop app it is still talking to OpenAI
Build · September 19, 2026 · 1 publisher
- SageMaker's fallback instance can serve a quantized copy of the same model
Build · September 18, 2026 · 1 publisher
- A tile-size clamp in vLLM's Triton kernel gets Gemma 4 running on a Tesla T4
Build · September 18, 2026 · 1 publisher
- NVIDIA's AIPerf drops the Perf Analyzer stack for worker processes over ZMQ
Build · September 18, 2026 · 1 publisher
- Kiro and Claude Code both picked a TGI container that could not load Qwen3
Build · September 18, 2026 · 1 publisher
- Vals put Hy4 Preview first among open-weight models on code migration at $3.41 a test
Build · September 17, 2026 · 1 publisher
- llama.cpp's -ngl flag keeps a 9B model on a 6GB card by leaving 28 layers on the CPU
Build · September 17, 2026 · 1 publisher
- Pointing the Ray head and workers at vLLM's own image removes the numpy 2.0 ABI crash
Build · September 17, 2026 · 1 publisher
- Thought-anchor scores change when a second model resamples the same trace
Build · September 16, 2026 · 1 publisher
- NVIDIA's 3.7x Vera Rubin figure comes from Qwen3-VL on vLLM and Dynamo
Build · September 16, 2026 · 1 publisher
- One slash in a Host header moves the path Starlette's middleware checks
Build · September 16, 2026 · 1 publisher
- Red Hat clocks sandbox isolation at under 5 percent of an agent request's latency
Product · September 15, 2026 · 1 publisher
- Whoever implements the server half of Responses picks your retrieval and veto defaults
Build · September 15, 2026 · 1 publisher
- Rubin shows 7x tokens per megawatt and over 2x profit per gigawatt versus Blackwell, SemiAnalysis says
Invest · September 14, 2026 · 1 publisher
- Red Hat traces the inference bill to 140GB of memory reads per token
Product · September 14, 2026 · 1 publisher
- Mixing self-hosted Qwen with Bedrock Claude costs you telemetry, not a rewrite
Build · August 14, 2026 · 1 publisher
- A backend that detects the AMD GPU can still leave operations on the CPU
Build · September 12, 2026 · 1 publisher
- 6,935 exposed Ollama servers answered an internet scan without asking for credentials
Security · September 11, 2026 · 1 publisher
- Renaming FastMCP to MCPServer closes stdio before an unpinned server finishes its handshake
Build · September 11, 2026 · 1 publisher
- Overnight laptop runs took over most of one Rust developer's Opus coding work
Build · September 10, 2026 · 1 publisher
- Routing on the prompt's first tokens cut AWS's median time to first token by up to 77 percent
Build · September 10, 2026 · 1 publisher
- NVIDIA's 2.5x concurrency figure rests on a 64K prompt reused 76 percent of the time
Build · September 10, 2026 · 1 publisher
- Why LLM output is hard to reproduce: it's not just concurrency and floating point
Science · September 10, 2026 · 1 publisher
- NVFP4 squeezes Qwen3.8's 2.4 trillion weights onto eight B300s at 150 GB a GPU
Build · September 9, 2026 · 1 publisher
- Judge model choice swings AI-Infra-Guard's false positive rate fifteenfold
Security · September 9, 2026 · 1 publisher