Skip to content

benchmark

WildGuard

Open safety moderation resource for classifying harmful prompts, harmful responses and refusals in language model interactions.

Current clusters