build1 publisher
Misbehaviour scores and past failure rates predict fine-tuning misalignment before training
AlignmentForecastBench researchers forecast fine-tuning misalignment above chance across 17 models and 32 datasets, using a data score and past failure rates. Frontier LLMs reading only the raw data barely beat chance, so the screen depends on a record of past fine-tuning failures.
Publishers:lesswrong.com
Reality
- Evidence45
- Adoption
- Insufficient
- Hype gap+10
- Incentives
- Insufficient
- Confidence40