Skip to content

Topic

LLM guardrail models

Small classifiers placed around a language model to score incoming or outgoing text for attacks, jailbreaks or policy violations before it is acted on.

Current clusters