VIDRAFT says POCKET-Darwin-180B, a 111 GB 4-bit GGUF build of its 180B mixture-of-experts model, runs in llama.cpp on about $1,400 of consumer hardware. The accuracy evidence so far is one MMLU-Pro comparison that VIDRAFT reports itself.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+45
- Incentives60
- Confidence35
VIDRAFT's prefix-invariance test reports causal leakage in Nemotron-H-8B and Zamba2-1.2B, starting at chunk sizes of 128 and 256. Mask inspection caught none of 192 injected faults.
Reality
- Evidence24
- Adoption8
- Hype gap+58
- Incentives86
- Confidence41
The $2,000 pot is noise. The scoring null, 20,000 random traders per asset re-drawn daily on the path that actually happened, is the part other leaderboards should copy.
Reality
- Evidence46
- Adoption12
- Hype gap+14
- Incentives68
- Confidence50
FINAL-Bench moved a fixed LightGBM baseline by 0.211 AUROC purely by changing how hERG data was divided. That is roughly eight times the spread across random seeds.
Reality
- Evidence55
- Adoption14
- Hype gap+12
- Incentives68
- Confidence48
VIDRAFT and FINAL-Bench opened a public leaderboard for AI-designed PfDHODH inhibitors. The instructive part is the fourteen defects they found in their own scoring system first.
Reality
- Evidence38
- Adoption17
- Hype gap+12
- Incentives70
- Confidence42