ModelRefs / MMLU Methodology — Methodology
MMLU Methodology — Methodology
MMLU evaluates broad multi-domain academic knowledge via 4-way multiple choice across 57 subjects.
Overview
What it measures: Breadth of factual and conceptual knowledge from elementary through professional levels.
How it works
- 57 subject categories ranging from elementary math to professional law.
- Each item is a 4-way multiple choice question.
- Models are scored on exact-match accuracy across ~14K questions.
- Typically evaluated 0-shot or 5-shot.
Strengths
- Wide subject coverage
- Cheap to evaluate
- Well-established baseline
Limitations
- Saturated by frontier models (>90%)
- Multiple-choice masks reasoning weaknesses
- Contamination risk in pretraining
Best use cases
- General-knowledge baselining
- Cross-model comparison at parity
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to MMLU Methodology — Methodology.
Frequently asked questions
What does MMLU measure?
Breadth of factual and conceptual knowledge from elementary through professional levels.
What are its main limitations?
Saturated by frontier models (>90%) Multiple-choice masks reasoning weaknesses Contamination risk in pretraining
When should I use this benchmark?
General-knowledge baselining Cross-model comparison at parity