ModelRefs / Best Reasoning Models

Best Reasoning Models

Top AI models ranked by reasoning benchmarks like MMLU, GPQA and ARC.

Overview

Top AI models ranked by reasoning benchmarks like MMLU, GPQA and ARC.

How this ranking is produced

30 models in the ModelRefs catalogue carry qualifying benchmark evidence for this category. The ten highest-scoring are listed below.

Scores below are a heuristic over the benchmark evidence ModelRefs holds for each model, not a guarantee of real-world performance. A model ranks only where it has qualifying benchmark results, so a capable model with thin evidence can rank low or be absent. Each entry states the benchmarks behind its score — read those before acting on the order.

Ranked models

  1. #1 Llama 3.1 405B

    Score 89 out of 100. Reasoning benchmark average of 88.6.

  2. #2 Gemini 1.5 Pro

    Score 86 out of 100. Reasoning benchmark average of 85.9.

  3. #3 GPT-5

    Score 86 out of 100. Reasoning benchmark average of 85.7.

  4. #4 Llama 3.1 70B

    Score 84 out of 100. Reasoning benchmark average of 83.6.

  5. #5 o3

    Score 83 out of 100. Reasoning benchmark average of 82.8.

  6. #6 Gemini 2.5 Flash

    Score 83 out of 100. Reasoning benchmark average of 82.8.

  7. #7 GPT-5 Mini

    Score 82 out of 100. Reasoning benchmark average of 82.3.

  8. #8 Claude Opus 4

    Score 80 out of 100. Reasoning benchmark average of 79.6.

  9. #9 Nemotron-4 340B

    Score 79 out of 100. Reasoning benchmark average of 78.7.

  10. #10 o4 Mini

    Score 78 out of 100. Reasoning benchmark average of 77.6.

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Best Reasoning Models.