Aleph Alpha released Kolibri, an Apache 2.0 German-English model that activates 3.46B of its 78.1B parameters per token. Each token costs about as much compute as a small model, yet a team hosting it in Europe still has to fit every expert in memory.
Perspective Coverage
5 publishers
- Builder
- Builder 49%
- Operator
- Operator 36%
- Investor
- Investor 15%
Reality
- Evidence70
- Adoption
- Insufficient
- Hype gap+15
- Incentives65
- Confidence68
Ollama since 0.34.4 lets Gemma 4 skip the requested JSON schema, returning bare text with HTTP 200 in 8 of 24 test calls. Until the open fix ships in a release, structured output on local thinking models needs a shape check in the client.
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap+10
- Incentives20
- Confidence58
Repacking Google's quantization-aware-trained Gemma 4 weights without re-rounding fits the 12B model on one TPU v5e chip at bf16-level accuracy. Adopting it means carrying patches to vLLM's TPU backend and working inside a 9,728-token KV cache.
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+15
- Incentives35
- Confidence55
Google's Cloud Run instances, priced from $5.70 a month, kept an open-source AI agent running with its memory intact in a developer's Preview test. For an agent that mostly waits, it trades a VM's patching for limits the code must be written around.
Reality
- Evidence40
- Adoption8
- Hype gap+15
- Incentives40
- Confidence38
Kaggle benchmark results posted on dev.to report that most of 30 vision models reading 336 synthetic Grafana-style panels found the peak but misread the clock. Copilot incident timelines drafted from screenshots need their start times checked by hand.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+5
- Incentives30
- Confidence45
SageMaker costs 1.40x a plain EC2 instance per hour to serve one Gemma 4 vLLM build on the same T4 or L4 GPU, a benchmark on dev.to finds. With decode speed matched within 2%, the extra 40% goes to the managed layer around the GPU.
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+5
- Incentives
- Insufficient
- Confidence60
Gemma 4 decodes on SageMaker's smallest GPU, a T4, at 0.8x an L4's speed with matching outputs from a patched vLLM image, a dev.to benchmark reports. The T4 is the cheaper choice per token only when it rents for under 80% of the L4's hourly rate.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+15
- Incentives35
- Confidence40
Repacking Gemma 4's QAT weights with 4-bit embedding tables made decode up to 1.39x faster on one SageMaker L4, according to a dev.to benchmark series. For teams serving Gemma 4 on vLLM, how the weights are stored becomes a setting to measure alongside model size.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+10
- Incentives
- Insufficient
- Confidence50
Developer xbill9 rebuilt Google's QAT Gemma 4 26B as int4 and fit it on one TPU v6e with 53,888 tokens of KV cache, against 3,456 for RedHat's FP8 build. Throughput nearly doubles, with accuracy checked on one classification suite.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+15
- Incentives40
- Confidence55
Gemma 4 E2B's 4-bit QAT checkpoint decodes 2.05x faster than bf16 on one SageMaker L4, according to a dev.to benchmark. The swap also frees 18% of GPU memory for a 20% larger KV cache, though it runs only through the vLLM container because JumpStart lists no QAT build.
Reality
- Evidence58
- Adoption
- Insufficient
- Hype gap+8
- Incentives
- Insufficient
- Confidence55
A single AMD Developer Cloud droplet at $1.99 an hour was inventoried entirely over twelve tag-scoped MCP tools, and two of the readings came back wrong until the output was parsed instead of a status code.
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap+12
- Incentives
- Insufficient
- Confidence55
NVIDIA's SWE-Serve scores the same 627 patches twice on 19 SGLang tasks, once with the live-serving tests and once without. The pass rate falls from 69.4% to 45.9%, and 242 of the 276 live tests came from SGLang itself.
Reality
- Evidence58
- Adoption30
- Hype gap−10
- Incentives70
- Confidence60
Red Hat clocks the same 20-call agent task at roughly 45 seconds on a slow backend and about 13 on a fast one. Model choice for agents is turning into a per-call latency budget, with capability as one input.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+10
- Incentives50
- Confidence55
Amazon says the Apache 2.0 model family is generally available on Bedrock's next-generation inference engine in eusc-de-east-1, reached through an OpenAI-compatible endpoint and run by staff who reside in the EU.
Reality
- Evidence34
- Adoption
- Insufficient
- Hype gap+15
- Incentives85
- Confidence58
On Alphabet's Q1 2026 call Sundar Pichai put cloud backlog above $460 billion, close to six years of cloud revenue at the current rate, while giving Gemini Enterprise's paid user base as a growth rate alone.
Reality
- Evidence55
- Adoption70
- Hype gap+25
- Incentives88
- Confidence62
Cisco's Splunk unit is repositioning around telemetry it expects agents to generate faster than people can read, with a new fabric that queries Snowflake, Databricks and object stores in place. Pricing was not disclosed.
Reality
- Evidence30
- Adoption15
- Hype gap+35
- Incentives85
- Confidence45
A Linux guest under Apple's container CLI cannot reach Metal. So the setup that works keeps the app in a Debian machine and calls Ollama on the Mac at 192.168.64.1. One run there measured 45.2 tok/s.
Reality
- Evidence58
- Adoption15
- Hype gap−5
- Incentives28
- Confidence55
VMware Private AI Cloud stacks inference, deny-by-default agent sandboxes and token metering on Cloud Foundation 9, so the sovereignty pitch now arrives attached to an infrastructure contract platform teams already signed.
Reality
- Evidence32
- Adoption28
- Hype gap+38
- Incentives84
- Confidence52
The Information reports a price that works out to 86 times Hugging Face's revenue, which means what Nvidia bought is the host of three million open-weight models, and every pipeline that fetches at build time has a new upstream owner.
Reality
- Evidence38
- Adoption46
- Hype gap+32
- Incentives68
- Confidence34
One hand-written JAX port runs on TPU v5e, v6e and a Turing T4G, but the quantized fast path stays on TPU and the compute dtype has to be read off the live device, because neither leak ever shows up as red in a log.
Reality
- Evidence54
- Adoption14
- Hype gap+8
- Incentives32
- Confidence45