Skip to content

other

Safety for Whom? Boundary-Aware Self-Distillation for Controlled LLM Safety Refusal

Research paper proposing that a deployment's refusal set be treated as a subset of a topic, with training data and evaluation built around the boundary between refusable and answerable prompts.

Known aliases

  • Safety for Whom?

Relationships

No evidence-backed relationships are recorded.

Current clusters