Salvatore Sanfilippo's open-source ds4 engine runs a short list of large open-weight models locally by compressing their routed experts to about two bits. Even compressed, the supported builds need high-memory Macs or GPU systems that most people do not already own.
Reality
- Evidence55
- Adoption30
- Hype gap+20
- Incentives
- Insufficient
- Confidence60
Six LLMs on a 12-task Kaggle security benchmark all caught SQL injection, hardcoded keys and pickle RCE, but Gemini 3.7 Flash missed a path traversal. With one scenario per flaw class, the run shows which textbook patterns the models know and says little about trusting one to review real code.
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+40
- Incentives40
- Confidence35
CTGT's 76-prompt audit puts an anonymous OpenRouter endpoint at 6.5 on the usual censorship probes and 89.5 on Chinese domestic legitimacy. The prompt list is the benchmark.
Reality
- Evidence55
- Adoption30
- Hype gap+12
- Incentives65
- Confidence52
A dev.to post scores ten coding models on five real tasks and divides by price. The method is cheap to copy; the vendor plumbing it recommends deserves more scrutiny than the arithmetic.
Reality
- Evidence20
- Adoption12
- Hype gap+45
- Incentives70
- Confidence55
IBM Research ran self-mined guidelines across eight models on AppWorld. One model gained 16.1 points for 5 percent more tokens; another gained nothing at all.
Reality
- Evidence52
- Adoption
- Insufficient
- Hype gap+15
- Incentives70
- Confidence45
Self-propagating payloads did move between agents through editable soul files, but one inoculation paragraph held against 150-plus optimized strains, and nothing propagated in the wild.
Reality
- Evidence66
- Adoption14
- Hype gap+18
- Incentives60
- Confidence55
Z.ai claims frontier agentic-coding scores at about 750B parameters, a third of Kimi K3, from extended post-training on the GLM-5.2 base. Open weights are promised in two weeks.
Publishers:interconnects.ai
Reality
- Evidence32
- Adoption24
- Hype gap+28
- Incentives68
- Confidence38
Z.ai's August 14 post claims post-training gains for coding agents, but the company's release notes still stop at GLM-5.1 and there is no API endpoint, model identifier or weight download.
Reality
- Evidence42
- Adoption18
- Hype gap+38
- Incentives68
- Confidence46