Skip to content

Topic

LLM benchmark design

The construction of evaluation suites for language models, covering dataset provenance, gold-label auditing, prompt isolation, scoring rules and how results are versioned.

Current clusters