ModelRefs / τ-bench Leaderboard — AI Model Scores
τ-bench Leaderboard — AI Model Scores
Tool-augmented agent benchmark on retail & airline customer-service tasks. Current leaders, methodology, and citation sources for τ-bench.
Overview
Tool-augmented agent benchmark on retail & airline customer-service tasks.
How it is measured: Pass^k success rate with policy adherence.
How this benchmark is scored
| Category | agents |
|---|---|
| Maximum score | 100 % success |
| Direction | Higher is better |
Primary source: https://github.com/sierra-research/tau-bench
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to τ-bench Leaderboard — AI Model Scores.