Default boosted trees came within half an AUC point of a 200-fit hyperparameter search on five of six public datasets, a dev.to benchmark found. Anyone reviewing an AI-generated modelling notebook can check the cheap baselines and the default trees first, and treat the tuning search as a cost it has to justify.
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+5
- Incentives
- Insufficient
- Confidence55
Fraudagent, a TigerGraph challenge entry, settles 20 fraud alerts in code whose build fails on any LLM import. Its verdicts match bit for bit with no API key, so the calibration bug its authors blame for clearing fraud lives in code a test can reach.
Reality
- Evidence45
- Adoption5
- Hype gap+10
- Incentives55
- Confidence45
A per-asset Venn-Abers audit of seven crypto models found ETH claiming 93.8% and delivering 45.5%, while XRP delivered 93.1% at 64.9% stated. Opposite errors need opposite corrections.
Reality
- Evidence54
- Adoption12
- Hype gap+6
- Incentives34
- Confidence47
FINAL-Bench moved a fixed LightGBM baseline by 0.211 AUROC purely by changing how hERG data was divided. That is roughly eight times the spread across random seeds.
Reality
- Evidence55
- Adoption14
- Hype gap+12
- Incentives68
- Confidence48
The round is small; the claim under it is testable. Synthefy says tables and time series need models built for them, and it has shipped a 30-million-parameter open model to be checked.
Reality
- Evidence28
- Adoption24
- Hype gap+42
- Incentives78
- Confidence33