Skip to content

Topic

Agent and tool-use evaluation

The practice of measuring whether language-model agents call the tools available to them correctly, including harness design, trial repetition and pass-rate scoring rules.

Current clusters