Amazon's new EKS managed addon reads Prometheus metrics from every model pod and sends each request to the one with room in its KV cache. The advertised 82% cut in first-token latency rests on a single 4.4-second baseline.
Reality
- Evidence46
- Adoption14
- Hype gap+38
- Incentives88
- Confidence52
TypeSafe's first model, Jev, answers structured questions with typed output and a confidence measure attached. The account of its launch says developers have to check those scores against real outcomes on their own data first.
Reality
- Evidence26
- Adoption7
- Hype gap+48
- Incentives80
- Confidence57
The one measured number behind NVIDIA's tokens-per-megawatt pitch comes from Lambda's Blackwell cluster, where power reclaimed from static provisioning ran three extra nodes. Amazon's Annapurna Labs and d-Matrix got a line each.
Reality
- Evidence33
- Adoption45
- Hype gap+44
- Incentives90
- Confidence60
Faraj Aalaei's Cognichip argues that semiconductor physics belongs inside the model, with agents left to orchestrate. The one speedup it has documented in detail is a 55-page spec handled in days against a baseline Cognichip supplied itself.
Reality
- Evidence33
- Adoption18
- Hype gap+42
- Incentives74
- Confidence57
Astra scored 98.6% on ARC-AGI-3 where its predecessor managed 7.8%. OpenAI is still rationing access while it scales capacity, and that tells a clerical-automation budget more than the benchmark does.
Reality
- Evidence36
- Adoption21
- Hype gap+41
- Incentives79
- Confidence47
Ritchie Vink is selling 2.0 as a boring upgrade. Mostly it is, but engine="auto" now resolves to streaming by default, so joins and group_bys stop carrying incidental row order, and pipelines that leaned on it will drift quietly.
Reality
- Evidence62
- Adoption20
- Hype gap+10
- Incentives78
- Confidence58
Gemini 3.8 Flash arrived three weeks after Google's last model, its cybersecurity sibling is invite-only, and Google says higher effort levels can spend more tokens. Standardising on a model name now has a shelf life.
Reality
- Evidence28
- Adoption26
- Hype gap+42
- Incentives84
- Confidence46
The CameraJet spots a gap between two teeth and fires rinse into it inside 100 milliseconds, which leaves the only proof that the machine learning worked sitting inside a phone app the toothbrush does not need.
Reality
- Evidence24
- Adoption9
- Hype gap+58
- Incentives78
- Confidence56
Mercury reports over 1,000 tokens per second per user, and Nemotron Diffusion claims 2 to 8 times autoregressive throughput. Both are vendor figures. The sampler that produces them charges for its speed in arithmetic.
Reality
- Evidence38
- Adoption42
- Hype gap+30
- Incentives62
- Confidence40
The zettaFLOPS figure is an estimate. The measured one, sitting in a GB300 example, is a 25 percent density gain, and it is paid for out of the safety margin.
Reality
- Evidence42
- Adoption15
- Hype gap+34
- Incentives78
- Confidence55
One SKU, one compute die, 1.2 TB/s of LPDDR5X, and benchmark wins over a 96-core EPYC. The comparison with AMD's Venice is being set on Nvidia's terms rather than on core counts.
Reality
- Evidence38
- Adoption12
- Hype gap+33
- Incentives78
- Confidence44
Qwen 3.8 27B ran on a MacBook Pro from a 17GB GGUF and spent 21 minutes on one SVG. Licensing and access stopped being the blocker; latency and KV cache budgeting became the job.
Reality
- Evidence48
- Adoption58
- Hype gap+18
- Incentives62
- Confidence52