Skip to content

Topic

LLM evaluation harnesses

Test suites that run fixed prompts against a language model and score the outputs, usually combining deterministic checks with human or model judgement.

Current clusters