Skip to content

Topic

Model guardrails

Techniques that constrain what a deployed language model outputs, including safety training, refusal tuning, system prompts and output filters.

Current clusters