Skip to content

Topic

AI control

Research programme on protocols that extract useful work from AI systems that may be adversarial, assuming the model itself cannot be trusted.

Current clusters