Skip to content

other

Model abliteration

An inference-time technique that projects away refusal-sensitive directions in a language model's activations, suppressing refusals without any weight training.

Current clusters