buildOne report1 publisher Epoch AI's InnovationEval credited GPT-5.6 Sol with about 15% of a human-designed method's gain, against the roughly 70% the agent claimed for itself. Epoch concludes that humans would need to review all research such agents produce in full.
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+45
- Incentives50
- Confidence55
buildConfirmed4 publishers Z.ai says the base model did not change between GLM-5.2 and GLM-5.3, so the coding jump and the doubled exploitation score come out of the same post-training run. Security teams inherit the second half.
Perspective Coverage
4 publishers
- Builder
- Builder 49%
- Operator
- Operator 32%
- Investor
- Investor 19%
Reality
- Evidence50
- Adoption30
- Hype gap+30
- Incentives70
- Confidence60
buildOne report1 publisher Zhipu says every gain in GLM 5.3 came from post-training on the same ~744B MoE it shipped as GLM 5.2. The largest claimed leap, in vulnerability discovery, is also the least quantified.
Reality
- Evidence25
- Adoption10
- Hype gap+35
- Incentives70
- Confidence35
buildOne report1 publisher ServiceNow's CoreAI team generates new agent training tasks, each with its own verifier, aimed at gaps a stronger teacher model can already solve. Adopting it means running a seedable copy of the environment and writing verifiers that accept any valid solution.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+10
- Incentives55
- Confidence40
buildConfirmed5 publishers Z.ai says every gain in GLM-5.3 came from post-training on an unchanged base. If that holds, refresh cadence for self-hosted weights is set by RL runs, not pretraining runs.
Perspective Coverage
5 publishers
- Builder
- Builder 58%
- Operator
- Operator 33%
- Investor
- Investor 9%
Reality
- Evidence40
- Adoption30
- Hype gap+35
- Incentives70
- Confidence55
buildOne report1 publisher An arXiv paper names the failure behavioral state decay and runs a side-car memory agent next to an unmodified action agent, reporting gains of 8.3 and 6.8 points of pass@1 on two long-horizon benchmarks.
Reality
- Evidence52
- Adoption
- Insufficient
- Hype gap+18
- Incentives60
- Confidence57
Harvey and Baseten report rubric pass rates rising from 23.3% to 62.4% across seven models once the data room sits in a Python REPL and sub-agents do the reading, while Claude Code with Opus-5 manages 24.6%.
Publishers:harvey.ai
Reality
- Evidence47
- Adoption17
- Hype gap+18
- Incentives84
- Confidence57
buildOne report1 publisher An empirical study of publicly released PostTrainBench trajectories finds the training strategy is fixed at step one and the whole remaining budget goes on local tweaks, with scaffolding and human hints improving execution and leaving that pattern intact.
Reality
- Evidence58
- Adoption20
- Hype gap+12
- Incentives35
- Confidence55
buildOne report1 publisher Microsoft reports Qwen3.5-9B going from 41.8% to 56.4% on SWE-bench Verified using 6K examples. The load-bearing change is who owns the agent loop, and what the service boundary costs.
Reality
- Evidence32
- Adoption18
- Hype gap+28
- Incentives66
- Confidence44