ModelRefs / GAIA Leaderboard — AI Model Scores

GAIA Leaderboard — AI Model Scores

General AI assistant benchmark — multi-tool real-world questions. Current leaders, methodology, and citation sources for GAIA.

Overview

General AI assistant benchmark — multi-tool real-world questions.

How it is measured: Exact-match across 3 difficulty levels.

How this benchmark is scored

Categoryagents
Maximum score100 % accuracy
DirectionHigher is better

Primary source: https://huggingface.co/gaia-benchmark

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to GAIA Leaderboard — AI Model Scores.