build1 publisher
Z-scoring the J-lens against base-model token frequency lifts hidden-word recovery to 0.805
The offset in a J-lens readout mostly tracks how often a token appears, and scaling it by variance, after plain subtraction failed, lifted hidden-word elicitation to 0.805 from 0.665 on Gemma-2-9B-it. The paired test over 20 words gives p of about 0.19.
Publishers:lesswrong.com
Reality
- Evidence46
- Adoption12
- Hype gap+10
- Incentives40
- Confidence56