Skip to content

Topic

LLM capability evaluation

The practice of testing what large language models can and cannot do, including safety-relevant tasks such as providing uplift on dangerous procedures.

Current clusters