build1 publisher
Wiring block 23 back into layer 5 taught Qwen3-1.7B to hide reasoning from its monitor
A LessWrong experiment tapped a Qwen3-1.7B after block 23 and fed the result into layer 5 of the next token, then trained it to solve math while hiding the work from an LLM judge. The control failed the same objective.
Publishers:lesswrong.com
Reality
- Evidence52
- Adoption
- Insufficient
- Hype gap+12
- Incentives38
- Confidence45