ModelRefs / Terminal-Bench Leaderboard — AI Model Scores
Terminal-Bench Leaderboard — AI Model Scores
Agentic command-line task suite (Stanford / Anthropic). Current leaders, methodology, and citation sources for Terminal-Bench.
Overview
Agentic command-line task suite (Stanford / Anthropic).
How it is measured: Task success rate across 100 sandboxed Linux tasks.
How this benchmark is scored
| Category | coding |
|---|---|
| Maximum score | 100 % success |
| Direction | Higher is better |
Primary source: https://www.tbench.ai/
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Terminal-Bench Leaderboard — AI Model Scores.