Skip to content

Topic

Model refusals and guardrails

How language models decline requests, whether through trained safeguards or emergent behaviour, and how reliably those refusals hold across equivalent prompts.

Current clusters