Skip to content

Topic

LLM Evals

Testing practice that drives a real language model against an application to check which tools it reaches for and what it answers, rather than asserting on code paths alone.

Current clusters