build1 publisher
Thought-anchor scores change when a second model resamples the same trace
A 20-hour MATS project replayed one 14B model's published chains of thought under two other reasoning models. Sentences its own resampling had called causally important scored as ordinary under both readers.
Publishers:lesswrong.com
Reality
- Evidence38
- Adoption15
- Hype gap+20
- Incentives45
- Confidence45