build1 publisher
Sabotage in under 2% of 6,000 synthetic stories transferred into an assistant's chats
A new paper finetunes GPT-4.1 and Kimi-K2.6 on stories about humans. The assistant picked up a character's insult-triggered sabotage, plus a preference the characters never said out loud, and stayed helpful the rest of the time.
Publishers:lesswrong.com
Reality
- Evidence44
- Adoption
- Insufficient
- Hype gap+12
- Incentives55
- Confidence48