ModelRefs / Best Open-Source Models

Best Open-Source Models

The most capable open-weight models you can self-host today.

Overview

The most capable open-weight models you can self-host today.

How this ranking is produced

24 models in the ModelRefs catalogue carry qualifying benchmark evidence for this category. The ten highest-scoring are listed below.

Scores below are a heuristic over the benchmark evidence ModelRefs holds for each model, not a guarantee of real-world performance. A model ranks only where it has qualifying benchmark results, so a capable model with thin evidence can rank low or be absent. Each entry states the benchmarks behind its score — read those before acting on the order.

Ranked models

  1. #1 Llama 3.1 405B

    Score 139 out of 100. Open-weight model (Llama 3.1 Community) with benchmark avg 88.6.

  2. #2 Phi-4

    Score 134 out of 100. Open-weight model (MIT) with benchmark avg 83.7.

  3. #3 Llama 3.1 70B

    Score 134 out of 100. Open-weight model (Llama 3.1 Community) with benchmark avg 83.6.

  4. #4 Nemotron-4 340B

    Score 129 out of 100. Open-weight model (NVIDIA Open Model) with benchmark avg 78.7.

  5. #5 DeepSeek V3

    Score 126 out of 100. Open-weight model (MIT) with benchmark avg 75.9.

  6. #6 Command R+

    Score 126 out of 100. Open-weight model (CC-BY-NC) with benchmark avg 75.7.

  7. #7 Llama 4 Scout

    Score 124 out of 100. Open-weight model (Llama 4 Community) with benchmark avg 74.3.

  8. #8 Qwen 2.5 72B

    Score 121 out of 100. Open-weight model (Qwen License) with benchmark avg 71.1.

  9. #9 Llama 3.1 8B

    Score 119 out of 100. Open-weight model (Llama 3.1 Community) with benchmark avg 69.4.

  10. #10 Mistral Nemo

    Score 118 out of 100. Open-weight model (Apache-2.0) with benchmark avg 68.0.

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Best Open-Source Models.