ModelRefs / τ-bench Leaderboard — AI Model Scores

τ-bench Leaderboard — AI Model Scores

Tool-augmented agent benchmark on retail & airline customer-service tasks. Current leaders, methodology, and citation sources for τ-bench.

Overview

Tool-augmented agent benchmark on retail & airline customer-service tasks.

How it is measured: Pass^k success rate with policy adherence.

How this benchmark is scored

Categoryagents
Maximum score100 % success
DirectionHigher is better

Primary source: https://github.com/sierra-research/tau-bench

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to τ-bench Leaderboard — AI Model Scores.