ModelRefs / GSM8K Leaderboard — AI Model Scores

GSM8K Leaderboard — AI Model Scores

Grade-school math word problems requiring multi-step arithmetic reasoning. Current leaders, methodology, and citation sources for GSM8K.

Overview

Grade-school math word problems requiring multi-step arithmetic reasoning.

How it is measured: Chain-of-thought; exact-match on final numeric answer.

How this benchmark is scored

Categoryreasoning
Maximum score100 % accuracy
DirectionHigher is better

Primary source: https://arxiv.org/abs/2110.14168

Published results

ModelScoreRun dateSource
GPT-597.52026-05-01Aggregated public reports
DeepSeek R195.52026-05-01Aggregated public reports
GPT-5 Mini922026-05-01Aggregated public reports
Mistral Large 2882026-05-01Aggregated public reports

Each score reflects the protocol and date of its own source run. Results from different harnesses are not directly comparable.

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to GSM8K Leaderboard — AI Model Scores.