ModelRefs / ∞Bench Leaderboard — AI Model Scores

∞Bench Leaderboard — AI Model Scores

Tasks exceeding 100K tokens spanning code, math, and retrieval. Current leaders, methodology, and citation sources for ∞Bench.

Overview

Tasks exceeding 100K tokens spanning code, math, and retrieval.

How it is measured: Mean score across 12 tasks.

How this benchmark is scored

Categoryreasoning
Maximum score100 score
DirectionHigher is better

Primary source: https://github.com/OpenBMB/InfiniteBench

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to ∞Bench Leaderboard — AI Model Scores.