Skip to content

benchmark

Hyper-τ-bench

Open-source benchmark from Sierra that measures how well an AI developer agent can build a working customer service agent from a simulated business's documents, transcripts, API and codebase, then scores the resulting agent on unseen conversations.

Known aliases

  • Hyper-tau-bench
  • Hyper-τ-bench benchmark

Relationships

No evidence-backed relationships are recorded.

Current clusters