ModelRefs / LongBench v2 Leaderboard — AI Model Scores
LongBench v2 Leaderboard — AI Model Scores
Realistic long-context tasks up to 2M tokens. Current leaders, methodology, and citation sources for LongBench v2.
Overview
Realistic long-context tasks up to 2M tokens.
How it is measured: Macro accuracy across 503 multiple-choice items.
How this benchmark is scored
| Category | reasoning |
|---|---|
| Maximum score | 100 % accuracy |
| Direction | Higher is better |
Primary source: https://longbench2.github.io/
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to LongBench v2 Leaderboard — AI Model Scores.