ModelRefs / MATH-500 Leaderboard — AI Model Scores
MATH-500 Leaderboard — AI Model Scores
500-problem competition-math subset used for reasoning model evals. Current leaders, methodology, and citation sources for MATH-500.
Overview
500-problem competition-math subset used for reasoning model evals.
How it is measured: Pass@1 on numeric/symbolic answers; CoT permitted.
How this benchmark is scored
| Category | reasoning |
|---|---|
| Maximum score | 100 % accuracy |
| Direction | Higher is better |
Primary source: https://github.com/openai/prm800k
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to MATH-500 Leaderboard — AI Model Scores.