NASA and IBM's open lunar model ingests 11 instrument modalities across a 20,000-to-1 resolution range. USRA's account of how those observations were aligned is the part that explains the result.
Reality
- Evidence60
- Adoption15
- Hype gap+10
- Incentives55
- Confidence65
IBM Research's ALTK-Evolve distils an agent's own trajectories into scored guidelines and injects the top five at inference time, and a companion post puts a number on the reliability an average success rate hides.
Reality
- Evidence45
- Adoption20
- Hype gap+15
- Incentives85
- Confidence55
A from-scratch binding to COIN-OR's Clp reached CRAN on 15 September 2026, bringing back the re-solve-from-a-saved-basis loop that R lost when clpAPI was archived, on a solver its own author says is no longer the fastest.
Reality
- Evidence58
- Adoption18
- Hype gap−10
- Incentives45
- Confidence62
IBM Research calls the 24-point drop between passing once and passing five times the consistency gap. Independence would have predicted a 27% five-run rate, so the failures are clustering on particular tasks.
Reality
- Evidence45
- Adoption20
- Hype gap+12
- Incentives45
- Confidence48
MediaTek led the earlier $5.4m and Etna Labs the $19.1m seed, and the pitch competes for a chip team's engineering-hours budget rather than its tool budget, on a benchmark that still leaves about one answer in ten wrong.
Reality
- Evidence27
- Adoption18
- Hype gap+42
- Incentives78
- Confidence34
IBM Research ran self-mined guidelines across eight models on AppWorld. One model gained 16.1 points for 5 percent more tokens; another gained nothing at all.
Reality
- Evidence52
- Adoption
- Insufficient
- Hype gap+15
- Incentives70
- Confidence45