Repacking Google's quantization-aware-trained Gemma 4 weights without re-rounding fits the 12B model on one TPU v5e chip at bf16-level accuracy. Adopting it means carrying patches to vLLM's TPU backend and working inside a 9,728-token KV cache.
Reality
- Evidence60
- Adoption
- Insufficient
- Hype gap+15
- Incentives35
- Confidence55
The tool pairs directional ablation with an automatic parameter search. Pulling refusal training out of a model now takes a command line and a consumer graphics card, and the community has already published more than 5,000 such models.
Reality
- Evidence42
- Adoption55
- Hype gap+18
- Incentives60
- Confidence55
Red Hat clocks the same 20-call agent task at roughly 45 seconds on a slow backend and about 13 on a fast one. Model choice for agents is turning into a per-call latency budget, with capability as one input.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+10
- Incentives50
- Confidence55
ShadowPEFT is in Hugging Face PEFT's main branch as of a September 15th announcement, and it carries its own hidden state and can be detached as a smaller standalone model. Trying it means installing PEFT from source.
Reality
- Evidence58
- Adoption20
- Hype gap+10
- Incentives65
- Confidence57
Decode is memory-bandwidth bound. Eight concurrent 128k Llama 3 70B streams need 97.8 ms of HBM transfer for every token generated, and DeepSeek's 512-scalar latent is aimed squarely at those bytes.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+40
- Incentives50
- Confidence50
A position paper argues that CLT-based intervals dramatically understate uncertainty below a few hundred datapoints. The specialized benchmarks frontier teams build are already smaller than that before anyone slices them by task.
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap+10
- Incentives30
- Confidence58
An empirical study of publicly released PostTrainBench trajectories finds the training strategy is fixed at step one and the whole remaining budget goes on local tweaks, with scaffolding and human hints improving execution and leaving that pattern intact.
Reality
- Evidence58
- Adoption20
- Hype gap+12
- Incentives35
- Confidence55
A new diagnostic benchmark treats the execution layer as something to vary rather than a fixed backdrop. Its conclusion is that agent capability belongs to a model-harness pair.
Reality
- Evidence42
- Adoption10
- Hype gap+22
- Incentives55
- Confidence40
OmnisBench's author rebuilt his split from date-stamped LiveCodeBench problems. The cheap tier fell from 90% to 60%, and a 4,096-token cap had been marking the frontier model absent.
Reality
- Evidence42
- Adoption
- Insufficient
- Hype gap+12
- Incentives78
- Confidence48
The update adds a path selector and a two-tap convolution rather than layers, recovering most of the accuracy that tripling the drafter bought at 15.2% latency, by the vendor's own numbers.
Reality
- Evidence54
- Adoption66
- Hype gap+16
- Incentives74
- Confidence58