ModelRefs / RULER 128K Leaderboard — AI Model Scores

RULER 128K Leaderboard — AI Model Scores

Long-context evaluation across 13 synthetic tasks at 128K tokens. Current leaders, methodology, and citation sources for RULER 128K.

Overview

Long-context evaluation across 13 synthetic tasks at 128K tokens.

How it is measured: Mean accuracy across NIAH, multi-hop, aggregation, QA.

How this benchmark is scored

Categoryreasoning
Maximum score100 % accuracy
DirectionHigher is better

Primary source: https://github.com/hsiehjackson/RULER

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to RULER 128K Leaderboard — AI Model Scores.