Multiverse Computing ranks candidate prunings by the energy of an Ising glass, so a single GPU scored about 29 billion combinations in two days. The paper reports the older per-block heuristics do just as well at lighter compression.
Reality
- Evidence50
- Adoption
- Insufficient
- Hype gap+25
- Incentives75
- Confidence55
The tool pairs directional ablation with an automatic parameter search. Pulling refusal training out of a model now takes a command line and a consumer graphics card, and the community has already published more than 5,000 such models.
Reality
- Evidence42
- Adoption55
- Hype gap+18
- Incentives60
- Confidence55
Red Hat clocks the same 20-call agent task at roughly 45 seconds on a slow backend and about 13 on a fast one. Model choice for agents is turning into a per-call latency budget, with capability as one input.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+10
- Incentives50
- Confidence55
A 20-hour MATS project replayed one 14B model's published chains of thought under two other reasoning models. Sentences its own resampling had called causally important scored as ordinary under both readers.
Reality
- Evidence38
- Adoption15
- Hype gap+20
- Incentives45
- Confidence45
A position paper argues that CLT-based intervals dramatically understate uncertainty below a few hundred datapoints. The specialized benchmarks frontier teams build are already smaller than that before anyone slices them by task.
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap+10
- Incentives30
- Confidence58
Stanford's AI Index puts a 142-fold parameter cut and a more than 280-fold price cut behind one fixed MMLU threshold. That narrows where building your own still pays, and it lands in a state-law count that doubled in a year.
Publishers:hai.stanford.edu
Reality
- Evidence60
- Adoption68
- Hype gap+10
- Incentives40
- Confidence57
The same guide that tells you to run evals on every change now carries a deprecation notice with two dates. For anyone whose deploy gate creates eval runs, the earlier date is the one that bites.
Reality
- Evidence72
- Adoption
- Insufficient
- Hype gap−28
- Incentives58
- Confidence68
The Engram explainer sells a second scaling axis beyond mixture-of-experts. The published comparison holds parameters and per-token compute fixed, which makes it a swap.
Reality
- Evidence52
- Adoption14
- Hype gap+32
- Incentives62
- Confidence55
A dev.to writeup grades a bug-fixing agent on its tool calls rather than its patch. Thirty documented runs cost roughly a dollar; the expensive input is the hand-built dataset behind them.
Reality
- Evidence34
- Adoption11
- Hype gap−8
- Incentives24
- Confidence41
A new diagnostic benchmark treats the execution layer as something to vary rather than a fixed backdrop. Its conclusion is that agent capability belongs to a model-harness pair.
Reality
- Evidence42
- Adoption10
- Hype gap+22
- Incentives55
- Confidence40
The acc and acc_norm split in lm-eval-harness can move in opposite directions on one checkpoint. Pick the metric before you train, and say which one you picked.
Reality
- Evidence52
- Adoption
- Insufficient
- Hype gap+14
- Incentives18
- Confidence46
Per-token prices fell about 200x since GPT-4's launch while US enterprise AI spend tripled to $37 billion. Tokenizer variance, reasoning tokens and tier discounts are where the bill diverges.
Publishers:hexaware.com
Reality
- Evidence55
- Adoption58
- Hype gap+15
- Incentives68
- Confidence45
Jinho Jang's Qwen3.8-27B-CRACK-GGUF packages abliterated multimodal weights with seven quantizations and a vision projector. The control point is now your inference hosts, not a vendor contract.
Reality
- Evidence42
- Adoption12
- Hype gap+18
- Incentives58
- Confidence55