ModelRefs / AIME 2024 Leaderboard — AI Model Scores
AIME 2024 Leaderboard — AI Model Scores
American Invitational Math Examination — 15-problem olympiad set. Current leaders, methodology, and citation sources for AIME 2024.
Overview
American Invitational Math Examination — 15-problem olympiad set.
How it is measured: Pass@1 over the 2024 paper; integer answers in [0,999].
What this benchmark measures
- competition mathematics
- multi-step symbolic reasoning
Relevant to:
- hard mathematical reasoning screening
- reasoning-model evaluation portfolios
Failure modes it exercises:
- incorrect exact-answer reasoning on bounded math problems
Method and its limits
Exact-match scoring on explicitly identified 2024 AIME I and/or II integer-answer problems; every report must state the form, prompt, sampling, and aggregation protocol.
- The two forms contain only 30 problems in total, so each item materially changes percentage accuracy.
- Reports differ on whether they use one or both forms, pass@1, majority vote, or repeated samples.
Dataset
- Dataset
- 2024 American Invitational Mathematics Examination
- Type
- high-school invitational competition mathematics problems
- Freshness
- unknown
AIME I and II each contain 15 problems with integer answers from 000 through 999; ModelRefs does not infer a benchmark freshness cutoff.
How to read this score
Report solved items and uncertainty, not only a percentage; compare scores only when form and sampling protocols match.
Similarity to real tasks: Low — AIME problems are rigorous but small, contest-style, exact-answer tasks rather than production mathematics workflows.
Data contamination risk: Unknown — The problems and solutions are public, and no model-specific training-overlap assessment is registered.
Benchmark gaming risk: Unknown — Small public sets are vulnerable to prompt, sampling, and benchmark-specific optimization effects.
What you still need to test yourself
- Test representative domain mathematics, proofs, tool use, numerical verification, and error recovery.
- Use hidden or newly authored problems to assess contamination and prompt sensitivity.
This benchmark supports decisions about:
- Screen exact-answer mathematical reasoning under a pinned 2024 form and protocol.
- Inspect problem-level errors before selecting a model for further math evaluation.
Limitations
- A small contest set does not establish broad mathematical competence or production reliability.
- A percentage without the exact form and sampling protocol is not comparable.
Dataset freshness is not fully confirmed from available source coverage. Contest and publication dates are not treated as model-training cutoffs.
Sources
- 2024 AIME I problems and answer key Art of Problem Solving, crediting the Mathematical Association of America · accessed 2026-06-29
- Model Card Addendum: Claude 3.5 Haiku and Upgraded Claude 3.5 Sonnet Anthropic · accessed 2026-06-29
How this benchmark is scored
| Category | reasoning |
|---|---|
| Maximum score | 100 % accuracy |
| Direction | Higher is better |
Primary source: https://artofproblemsolving.com/
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to AIME 2024 Leaderboard — AI Model Scores.