Skip to content

benchmark

Berkeley Function-Calling Leaderboard

Tool-use benchmark cited for the ~89% score of llama3-groq-tool-use:8b, which the author argues measures single well-formed function calls rather than multi-step agentic reliability.

Known aliases

  • Berkeley Function Calling Leaderboard v4
  • BFCL

Relationships

No evidence-backed relationships are recorded.

Current stories