Skip to content

Topic

LLM refusal and abstention

Work on when a language model declines to answer, covering safety refusal of harmful requests, abstention on questions with no valid answer, and the training that produces either behaviour.

Current clusters