Skip to content

benchmark

tau-bench

Benchmark testing AI agents on realistic tool use and multi-turn dialogue with simulated users, across domains like airline, retail, and banking.

Known aliases

  • tau3-Bench Banking
  • tau-airline
  • tau-retail
  • Tool-Agent-User Interaction Benchmark
  • τ³-Bench
  • τ-bench

Relationships

No evidence-backed relationships are recorded.

Current stories

build6 publishers

Spark 1.3's index jump lands on the three tests that carry half the score

Meta's fourth Muse Spark in five months gains four points on Artificial Analysis' Intelligence Index, almost entirely in agentic tests, while two scores fall and the tier Meta benchmarked against rivals is still a limited preview.

Perspective Coverage

6 publishers
Builder
Builder 52%
Operator
Operator 26%
Investor
Investor 22%

Reality

Evidence68
Adoption25
Hype gap+30
Incentives65
Confidence70