Skip to content

Topic

Safety alignment

The post-training work that makes a language model decline certain requests, together with research into how durable those refusals are once weights are public.

Current clusters