ModelRefs / Agents Benchmarks — Top AI Models

Agents Benchmarks — Top AI Models

Tool-use, function-calling, web/computer-use, and multi-step task execution. Measure end-to-end agentic task completion under real environments.

Overview

Tool-use, function-calling, web/computer-use, and multi-step task execution.

What this category is for: Measure end-to-end agentic task completion under real environments.

Benchmarks in this category

  • τ-bench — Tool-augmented agent benchmark on retail & airline customer-service tasks.
  • BFCL v3 — Berkeley Function-Calling Leaderboard, parallel + multi-turn tool use.
  • WebArena — Realistic web agent tasks across 5 self-hosted sites.
  • OSWorld — Computer-use agents across real OS apps (Ubuntu/Windows/macOS).
  • GAIA — General AI assistant benchmark — multi-tool real-world questions.

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Agents Benchmarks — Top AI Models.