security1 publisher
Heretic automatically strips refusals from many open-weight language models
The tool pairs directional ablation with an automatic parameter search. Pulling refusal training out of a model now takes a command line and a consumer graphics card, and the community has already published more than 5,000 such models.
Publishers:github.com
Reality
- Evidence42
- Adoption55
- Hype gap+18
- Incentives60
- Confidence55