ModelRefs / AIME 2025 Leaderboard — AI Model Scores

AIME 2025 Leaderboard — AI Model Scores

The two 2025 AIME competition-mathematics forms used for exact-answer reasoning evaluation. Current leaders, methodology, and citation sources for AIME 2025.

Overview

The two 2025 AIME competition-mathematics forms used for exact-answer reasoning evaluation.

How it is measured: Exact-answer accuracy on AIME I, AIME II, or both; form coverage, tools, sampling, answer extraction, and aggregation must be reported per run.

What this benchmark measures

  • competition mathematics
  • multi-step exact-answer reasoning

Relevant to:

  • recent mathematical reasoning screening
  • reasoning-model evaluation portfolios

Failure modes it exercises:

  • incorrect exact-answer reasoning on a recent bounded problem set

Method and its limits

Exact-match evaluation on explicitly identified 2025 AIME I and/or II problems; reports must state form, prompt, sampling, answer extraction, and aggregation.

  • The combined set is only 30 problems and early reports may cover only AIME I.
  • Release recency reduces some exposure pathways but does not prove absence from training, browsing, tools, or benchmark tuning.
  • Provider results are not automatically comparable: disclosures vary across form coverage, tool access, reasoning effort, sampling, answer extraction, and parallel-selection or majority-vote aggregation.

Dataset

Dataset
2025 American Invitational Mathematics Examination
Type
high-school invitational competition mathematics problems
Freshness
aging

Coverage is established for the complete 2025 competition artifact: AIME I was administered February 6 and AIME II February 12, 2025, with 15 integer-answer problems per form (30 distinct problems when both forms are combined). Reports must still identify whether they evaluated AIME I, AIME II, or both; February 12 records completion of the two-form administration, not a model-training cutoff, contamination clearance, or model-score evaluation date.

How to read this score

State the form and denominator, show problem-level outcomes, and avoid comparing one-form results with combined-form results.

Similarity to real tasks: Low — Established from the official competition format: AIME is a proctored, three-hour invitational examination with 15 difficult pre-calculus problems per form and an exact integer answer from 000 through 999 for each problem. The 2025 I and II forms therefore resemble a bounded human contest-mathematics task and provide a useful check of multi-step symbolic reasoning under exact-answer scoring. They do not reproduce most professional mathematics work: no proof or derivation is graded, and the task does not require requirements discovery, current evidence, domain software, numerical validation, peer review, uncertainty reporting, or a reusable technical deliverable. Real-task similarity is supported but low rather than unavailable.

Data contamination risk: High — The complete 30-problem set and answer material became publicly inspectable after the February 2025 administrations, and the small fixed set is now repeatedly used in model reports. This creates high exposure and memorization risk for post-release models and tool-enabled evaluations. The classification describes benchmark-level exposure; no model-specific training-overlap, browsing, tool-access, or post-release contamination audit is inferred.

Benchmark gaming risk: High — A 30-item exact-answer set is highly sensitive to form selection, prompt and answer parsing, reasoning budget, repeated sampling, and aggregation. Anthropic's Claude 4 disclosure, for example, separates standard nucleus-sampled results from higher parallel-test-time-compute results selected by an internal scoring model. Every comparison must therefore pin the evaluated form, tools, sampling, extraction, and aggregation protocol.

What you still need to test yourself

  • Test hidden, domain-relevant, tool-assisted, proof, and numerical-verification tasks.
  • Assess repeated-run variance, prompt sensitivity, leakage pathways, latency, and cost.

This benchmark supports decisions about:

  • Screen recent contest-math performance under a pinned form and reproducible protocol.
  • Identify failure patterns for deeper internal math evaluation.

Limitations

  • Recency is not proof of uncontaminated evaluation.
  • Small-set accuracy does not establish broad or production mathematical reliability.

Dataset publication coverage is confirmed through the February 12, 2025 AIME II administration and is now aging. These official competition dates establish when the two 2025 forms existed, but do not establish model-training isolation, post-release exposure, or the freshness of any reported model score.

Sources

How this benchmark is scored

Categoryreasoning
Maximum score100 % accuracy
DirectionHigher is better

Primary source: https://maa.org/maa-invitational-competitions/

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to AIME 2025 Leaderboard — AI Model Scores.