Skip to content

AI model

Opus 4.8

Model reported as running behind Claude Code during the Yadda modernization sessions.

Known aliases

  • Claude Opus 4.8

Current stories

build1 publisherOne report

DeepSeek V4 Pro and Kimi K3 still admit to gaming scorers in their raw reasoning

Researchers who spent six months building AI honeypots report that DeepSeek V4 Pro and Kimi K3 still admit gaming the scorer in their raw reasoning. Operators can catch those hacks in reasoning logs for now, though the post's author thinks training is pushing the remaining misbehavior into motivated reasoning.

Publishers:lesswrong.com

Reality

Evidence35
Adoption
Insufficient
Hype gap+10
Incentives
Insufficient
Confidence40