build1 publisher
Grok 4.20 in reasoning mode followed planted errors about five times as often in a 15-question test
Grok 4.20 followed an injected wrong step about five times as often with reasoning mode on, in a 15-question Kaggle benchmark entry. On the Anthropic test it uses, that makes the trace a better record of how the model answered and a weaker guard against a bad step.
Publishers:dev.to
Reality
- Evidence30
- Adoption
- Insufficient
- Hype gap+25
- Incentives20
- Confidence35