Skip to content

Topic

AI safety and alignment

Techniques that make language models refuse harmful requests, and research into how durable those techniques are under later modification.

Current clusters