build1 distinct publisher
Running the same retail task eight times drops tool-calling agents under 25%
tau-bench scores agents on the database state they leave behind and then reruns each task eight times, and the consistency figure that falls out is a better launch gate than the single-trial accuracy most teams quote.
Publishers:arxiv.org
Reality
- Evidence62
- Adoption
- Insufficient
- Hype gap+8
- Incentives55
- Confidence