ModelRefs / MATH-500 Leaderboard — AI Model Scores

MATH-500 Leaderboard — AI Model Scores

500-problem competition-math subset used for reasoning model evals. Current leaders, methodology, and citation sources for MATH-500.

Overview

500-problem competition-math subset used for reasoning model evals.

How it is measured: Pass@1 on numeric/symbolic answers; CoT permitted.

How this benchmark is scored

Categoryreasoning
Maximum score100 % accuracy
DirectionHigher is better

Primary source: https://github.com/openai/prm800k

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to MATH-500 Leaderboard — AI Model Scores.