build1 publisher
Ten human-labelled prompts calibrate the judges in this abliteration study
Abliteration needs no gradient updates, so a projection at inference time is a fair test of whether a data recipe actually diffused refusal behaviour. The judge you pick to score it changes the answer, and this protocol calibrates its judges on ten prompts.
Publishers:arxiv.org
Reality
- Evidence55
- Adoption
- Insufficient
- Hype gap+12
- Incentives30
- Confidence45