Skip to content

Topic

AI safety classifiers

Filters that inspect prompts and model output at runtime and block requests a developer judges dangerous, at the cost of some false positives on benign work.

Current clusters