ModelRefs / BIG-Bench Hard Leaderboard — AI Model Scores

BIG-Bench Hard Leaderboard — AI Model Scores

23 challenging tasks from BIG-Bench where prior LMs underperformed humans. Current leaders, methodology, and citation sources for BIG-Bench Hard.

Overview

23 challenging tasks from BIG-Bench where prior LMs underperformed humans.

How it is measured: 3-shot CoT; reported as macro-average accuracy.

How this benchmark is scored

Categoryreasoning
Maximum score100 % accuracy
DirectionHigher is better

Primary source: https://arxiv.org/abs/2210.09261

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to BIG-Bench Hard Leaderboard — AI Model Scores.