ModelRefs / MMLU Leaderboard — AI Model Scores
MMLU Leaderboard — AI Model Scores
Massive Multitask Language Understanding — 57 academic and professional subjects. Current leaders, methodology, and citation sources for MMLU.
Overview
Massive Multitask Language Understanding — 57 academic and professional subjects.
How it is measured: 5-shot multiple choice across 57 subjects; reported as overall accuracy.
What this benchmark measures
- academic knowledge
- multiple-choice reasoning
Relevant to:
- broad knowledge screening
- reasoning evaluation portfolios
- limited subject-level screening for clinical, legal, policy, coding, and investor workflow evaluation planning
Failure modes it exercises:
- knowledge and reasoning errors across covered subjects
Method and its limits
Five-shot multiple-choice accuracy across 57 academic and professional subjects.
- Aggregate accuracy can hide large subject-level differences.
- Multiple-choice performance does not establish generation, tool-use, or production reliability.
Dataset
- Dataset
- Massive Multitask Language Understanding
- Type
- academic and professional multiple-choice questions
- Freshness
- aging
The source describes 57 subjects and the author repository provides the original evaluation assets. The 2020 release date is a benchmark-release reference, not a current-world knowledge cutoff.
How to read this score
Overall accuracy is readable, but should be paired with subject-level results and task-specific evaluation.
Similarity to real tasks: Low — Established from the primary methodology: MMLU uses five-shot multiple-choice accuracy across 57 academic and professional subjects, including mathematics, history, computer science, medicine, law, and ethics. This provides a broad, controlled check of recalled knowledge and answer selection that can expose subject-level weaknesses. It does not reproduce end-to-end professional work: the model is not required to gather current evidence, resolve an ambiguous brief, use domain tools, apply jurisdictional or organizational context, create a reviewable deliverable, cite sources, communicate uncertainty, or remain reliable across a continuing workflow. Real-task similarity is therefore supported but low rather than unavailable.
Data contamination risk: High — MMLU is a public, long-running benchmark with downloadable evaluation material. ModelRefs cannot rule out training overlap for any specific model run, so current use should treat contamination exposure as high unless the score source documents decontamination.
Benchmark gaming risk: High — MMLU is widely used and saturated in many frontier-model reports; prompt, harness, and benchmark-specific optimization can materially affect interpretation.
What you still need to test yourself
- Evaluate current, domain-specific, open-ended, and production-format tasks.
- Measure reliability, calibration, safety, latency, cost, and deployment constraints separately.
- For clinical, legal, coding, payer-policy, and investor workflows, use current source-grounded cases with qualified reviewers, citation checks, jurisdiction or policy scope, consequential-error analysis, and explicit abstention.
This benchmark supports decisions about:
- Compare broad academic and professional task performance under a matched harness.
- Identify subject areas that warrant deeper task-specific evaluation.
Limitations
- Not a substitute for evaluation on current, domain-specific, or open-ended tasks.
- Does not measure cost, latency, safety, privacy, or deployment fit.
- Subject-level questions do not establish clinical judgment, legal research validity, coding correctness, coverage interpretation, disclosure compliance, or investor-communication readiness.
- Every model result still needs a canonical run record naming the model version, MMLU protocol, score source, methodology caveats, and limitations.
Sources re-reviewed 2026-07-11. MMLU is treated as an aging 2020 benchmark: the release date is known, but it is not a current-world knowledge cutoff and does not establish score freshness.
Sources
- Measuring Massive Multitask Language Understanding MMLU authors / ICLR 2021 · accessed 2026-07-11
- MMLU test and evaluation code MMLU authors · accessed 2026-07-11
How this benchmark is scored
| Category | reasoning |
|---|---|
| Maximum score | 100 % accuracy |
| Direction | Higher is better |
Primary source: https://arxiv.org/abs/2009.03300
Published results
| Model | Score | Run date | Source |
|---|---|---|---|
| GPT-5 | 89.2 | 2026-05-01 | Aggregated public reports |
| Claude Opus 4 | 88.7 | 2026-05-01 | Aggregated public reports |
| DeepSeek R1 | 84.1 | 2026-05-01 | Aggregated public reports |
| GPT-5 Mini | 82.4 | 2026-05-01 | Aggregated public reports |
| Mistral Large 2 | 78.4 | 2026-05-01 | Aggregated public reports |
| Llama 4 Scout | 76.8 | 2026-05-01 | Aggregated public reports |
| Command R+ | 75.7 | 2026-05-01 | Aggregated public reports |
Each score reflects the protocol and date of its own source run. Results from different harnesses are not directly comparable.
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to MMLU Leaderboard — AI Model Scores.