ModelRefs / MMLU Leaderboard — AI Model Scores

MMLU Leaderboard — AI Model Scores

Massive Multitask Language Understanding — 57 academic and professional subjects. Current leaders, methodology, and citation sources for MMLU.

Overview

Massive Multitask Language Understanding — 57 academic and professional subjects.

How it is measured: 5-shot multiple choice across 57 subjects; reported as overall accuracy.

What this benchmark measures

  • academic knowledge
  • multiple-choice reasoning

Relevant to:

  • broad knowledge screening
  • reasoning evaluation portfolios
  • limited subject-level screening for clinical, legal, policy, coding, and investor workflow evaluation planning

Failure modes it exercises:

  • knowledge and reasoning errors across covered subjects

Method and its limits

Five-shot multiple-choice accuracy across 57 academic and professional subjects.

  • Aggregate accuracy can hide large subject-level differences.
  • Multiple-choice performance does not establish generation, tool-use, or production reliability.

Dataset

Dataset
Massive Multitask Language Understanding
Type
academic and professional multiple-choice questions
Freshness
aging

The source describes 57 subjects and the author repository provides the original evaluation assets. The 2020 release date is a benchmark-release reference, not a current-world knowledge cutoff.

How to read this score

Overall accuracy is readable, but should be paired with subject-level results and task-specific evaluation.

Similarity to real tasks: Low — Established from the primary methodology: MMLU uses five-shot multiple-choice accuracy across 57 academic and professional subjects, including mathematics, history, computer science, medicine, law, and ethics. This provides a broad, controlled check of recalled knowledge and answer selection that can expose subject-level weaknesses. It does not reproduce end-to-end professional work: the model is not required to gather current evidence, resolve an ambiguous brief, use domain tools, apply jurisdictional or organizational context, create a reviewable deliverable, cite sources, communicate uncertainty, or remain reliable across a continuing workflow. Real-task similarity is therefore supported but low rather than unavailable.

Data contamination risk: High — MMLU is a public, long-running benchmark with downloadable evaluation material. ModelRefs cannot rule out training overlap for any specific model run, so current use should treat contamination exposure as high unless the score source documents decontamination.

Benchmark gaming risk: High — MMLU is widely used and saturated in many frontier-model reports; prompt, harness, and benchmark-specific optimization can materially affect interpretation.

What you still need to test yourself

  • Evaluate current, domain-specific, open-ended, and production-format tasks.
  • Measure reliability, calibration, safety, latency, cost, and deployment constraints separately.
  • For clinical, legal, coding, payer-policy, and investor workflows, use current source-grounded cases with qualified reviewers, citation checks, jurisdiction or policy scope, consequential-error analysis, and explicit abstention.

This benchmark supports decisions about:

  • Compare broad academic and professional task performance under a matched harness.
  • Identify subject areas that warrant deeper task-specific evaluation.

Limitations

  • Not a substitute for evaluation on current, domain-specific, or open-ended tasks.
  • Does not measure cost, latency, safety, privacy, or deployment fit.
  • Subject-level questions do not establish clinical judgment, legal research validity, coding correctness, coverage interpretation, disclosure compliance, or investor-communication readiness.
  • Every model result still needs a canonical run record naming the model version, MMLU protocol, score source, methodology caveats, and limitations.

Sources re-reviewed 2026-07-11. MMLU is treated as an aging 2020 benchmark: the release date is known, but it is not a current-world knowledge cutoff and does not establish score freshness.

Sources

How this benchmark is scored

Categoryreasoning
Maximum score100 % accuracy
DirectionHigher is better

Primary source: https://arxiv.org/abs/2009.03300

Published results

ModelScoreRun dateSource
GPT-589.22026-05-01Aggregated public reports
Claude Opus 488.72026-05-01Aggregated public reports
DeepSeek R184.12026-05-01Aggregated public reports
GPT-5 Mini82.42026-05-01Aggregated public reports
Mistral Large 278.42026-05-01Aggregated public reports
Llama 4 Scout76.82026-05-01Aggregated public reports
Command R+75.72026-05-01Aggregated public reports

Each score reflects the protocol and date of its own source run. Results from different harnesses are not directly comparable.

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to MMLU Leaderboard — AI Model Scores.