ModelRefs / Best Reasoning Models
Best Reasoning Models
Top AI models ranked by reasoning benchmarks like MMLU, GPQA and ARC.
Overview
Top AI models ranked by reasoning benchmarks like MMLU, GPQA and ARC.
How this ranking is produced
30 models in the ModelRefs catalogue carry qualifying benchmark evidence for this category. The ten highest-scoring are listed below.
Scores below are a heuristic over the benchmark evidence ModelRefs holds for each model, not a guarantee of real-world performance. A model ranks only where it has qualifying benchmark results, so a capable model with thin evidence can rank low or be absent. Each entry states the benchmarks behind its score — read those before acting on the order.
Ranked models
-
#1 Llama 3.1 405B
Score 89 out of 100. Reasoning benchmark average of 88.6.
-
#2 Gemini 1.5 Pro
Score 86 out of 100. Reasoning benchmark average of 85.9.
-
#3 GPT-5
Score 86 out of 100. Reasoning benchmark average of 85.7.
-
#4 Llama 3.1 70B
Score 84 out of 100. Reasoning benchmark average of 83.6.
-
#5 o3
Score 83 out of 100. Reasoning benchmark average of 82.8.
-
#6 Gemini 2.5 Flash
Score 83 out of 100. Reasoning benchmark average of 82.8.
-
#7 GPT-5 Mini
Score 82 out of 100. Reasoning benchmark average of 82.3.
-
#8 Claude Opus 4
Score 80 out of 100. Reasoning benchmark average of 79.6.
-
#9 Nemotron-4 340B
Score 79 out of 100. Reasoning benchmark average of 78.7.
-
#10 o4 Mini
Score 78 out of 100. Reasoning benchmark average of 77.6.
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Best Reasoning Models.