Skip to content

project

ThinkSafe

Safety-tuning approach that builds refusal training data by steering a target model into refusals and keeping the traces a guard model verifies as genuine.

Current clusters