Microsoft published its retrosynthesis model RetroChimera in Nature and released the weights. The blind test behind the headline claim graded individual disconnections, a narrower question than whether a whole route survives a lab.
Reality
- Evidence62
- Adoption28
- Hype gap+18
- Incentives68
- Confidence55
Tencent's Hunyuan Speech team put a swappable agent model behind a conversation model that re-decides every second whether to talk, and reports best-in-test timing on Full-Duplex-Bench v3 with a task-accuracy shortfall it describes only as slight.
Reality
- Evidence34
- Adoption12
- Hype gap+18
- Incentives62
- Confidence45
Alibaba's DAMO Academy put RADAR's training code and pretrained checkpoints under Apache 2.0 on September 18, a day after the Science paper. The 146 in the coverage counts radiological findings, cancers among them.
Reality
- Evidence70
- Adoption20
- Hype gap+20
- Incentives70
- Confidence62
VIDRAFT AI Research spread seven sequence mixers over 49 layers on a 7x7 Latin square so depth placement could not confound composition. The ablation that carries the finding ran on a 16-layer proxy at 700.9M parameters.
Reality
- Evidence48
- Adoption20
- Hype gap+15
- Incentives40
- Confidence55
Atlassian's own validator did the grading, on 25 app briefs, on the same workstation that trained the adapter. Anyone reusing the pass rate needs briefs that look like those 25.
Reality
- Evidence55
- Adoption15
- Hype gap+10
- Incentives65
- Confidence45
IBM research data scientist Gabby Nyirjesy told Lets Data Science that neighboring map patches went into training and test with no separation between them. The 22% gain over SwinV2-B holds under that one partition.
Reality
- Evidence70
- Adoption22
- Hype gap−10
- Incentives58
- Confidence65
The speed figure comes from a single dev.to write-up and describes generation, not delivery. The design decision underneath it is that frames and stereo audio leave in one pass, so the smallest thing you can retry is the whole clip.
Reality
- Evidence32
- Adoption34
- Hype gap+35
- Incentives80
- Confidence45
Six models from 0.9B to 375B, four of them pretrained on one shared token sequence, mean IFM's sparse-versus-dense efficiency claim can be tested by outsiders at a scale they can actually afford to re-run.
Reality
- Evidence44
- Adoption18
- Hype gap+16
- Incentives68
- Confidence52
The 260M NeoMME sends raw image patches and text down the same bidirectional path, with no vision tower and no causal decoder, and reports a 6 kB late-interaction index per page. The conditions behind both numbers are checkable.
Reality
- Evidence45
- Adoption18
- Hype gap+20
- Incentives75
- Confidence40
DeepSeek V4 Flash Vision Exp undercuts Gemini 3.7 Flash by 3.4x on input tokens and 5.7x on output, then spent 3,467 completion tokens and 30.5 seconds on an invoice both models audited correctly.
Reality
- Evidence56
- Adoption37
- Hype gap+26
- Incentives68
- Confidence55
Mercury reports over 1,000 tokens per second per user, and Nemotron Diffusion claims 2 to 8 times autoregressive throughput. Both are vendor figures. The sampler that produces them charges for its speed in arithmetic.
Reality
- Evidence38
- Adoption42
- Hype gap+30
- Incentives62
- Confidence40
The two-building region closes the data-at-rest question for Brazilian workloads, but the nearest sibling site is in Mexico, and the ten agent products Alibaba is marketing alongside it still carry no date.
Reality
- Evidence40
- Adoption34
- Hype gap+32
- Incentives78
- Confidence44
A week-long GLM-5.3-Flash preview ran on tens of thousands of domestic chips at per-token cost Z.ai says matches mainstream Nvidia parts, but with no vendor named and no power or throughput figures, the result speaks to inference sourcing and not to training.
Reality
- Evidence38
- Adoption46
- Hype gap+40
- Incentives78
- Confidence52
The model now writes its own training tasks and grading harnesses. That removes the bottleneck of hand-built tasks and replaces it with a harder one: rewards that cannot be gamed.
Reality
- Evidence38
- Adoption20
- Hype gap+24
- Incentives72
- Confidence54
Alibaba scheduled a 27-billion-parameter vision-language model for August 14, alongside an already-published 2.4-trillion-parameter MoE. The smaller file is the consequential one.
Reality
- Evidence30
- Adoption10
- Hype gap+35
- Incentives75
- Confidence32
Zhipu says GLM-5.3 edged Anthropic and OpenAI on one security benchmark. On the harder exploitation test the gap runs the other way, by 23.6 points.
Reality
- Evidence24
- Adoption18
- Hype gap+58
- Incentives79
- Confidence46