ModelRefs / Best Open Source AI Models
Best Open Source AI Models
Self-hostable open-weight models ranked by measured capability across reasoning, coding, and instruction following.
Overview
This page answers: Which open source model should I self-host?
How this ranking is produced
Recommendation engine phase-2.0.0, 16 benchmarks evaluated, aggregate freshness fresh. Computed from the canonical ModelRefs registry when this page was built on 2026-09-04, and recomputed on every deploy.
Each candidate below is a provisional fit signal under the stated constraints, not a guarantee, a certification, or a final ranking. Fit scores compare models against this use case's capability weights using current ModelRefs evidence; they are not accuracy rates, benchmark results, or production-readiness claims.
Ranked candidates
-
#1 Phi-4 — Microsoft
Fit score 63 out of 100.
- Strong cost efficiency (100/100).
- Strong multilingual (81/100).
- Open-weight (MIT).
- Caution: Weak hallucination resistance (3/100).
Confidence: 40% · Freshness: Stale 0% · Evidence: 67% · Benchmarks: 67% · Stability: 88%
-
#2 Llama 3.1 405B — Meta
Fit score 61 out of 100.
- Strong reasoning (89/100).
- Strong instruction following (89/100).
- Open-weight (Llama 3.1 Community).
- Caution: Limited benchmark coverage (1 scores).
Confidence: 40% · Freshness: Evaluation date not disclosed · Evidence: 33% · Benchmarks: 33% · Stability: 13%
-
#3 Llama 3.1 70B — Meta
Fit score 58 out of 100.
- Strong cost efficiency (86/100).
- Strong reasoning (84/100).
- Strong instruction following (84/100).
- Open-weight (Llama 3.1 Community).
- Caution: Limited benchmark coverage (1 scores).
Confidence: 40% · Freshness: Evaluation date not disclosed · Evidence: 33% · Benchmarks: 33% · Stability: 13%
-
#4 Mixtral 8x7B — Mistral AI
Fit score 57 out of 100.
- Strong cost efficiency (91/100).
- Strong reasoning (83/100).
- Strong instruction following (83/100).
- Open-weight (Apache-2.0).
- Caution: Limited benchmark coverage (1 scores).
Confidence: 40% · Freshness: Evaluation date not disclosed · Evidence: 33% · Benchmarks: 33% · Stability: 13%
-
#5 Nemotron-4 340B — NVIDIA
Fit score 54 out of 100.
- Strong cost efficiency (80/100).
- Strong reasoning (79/100).
- Strong instruction following (79/100).
- Open-weight (NVIDIA Open Model).
- Caution: Limited benchmark coverage (1 scores).
Confidence: 40% · Freshness: Evaluation date not disclosed · Evidence: 33% · Benchmarks: 33% · Stability: 13%
-
#6 Command R+ — Cohere
Fit score 52 out of 100.
- Strong reasoning (76/100).
- Strong instruction following (76/100).
- Open-weight (CC-BY-NC).
- Caution: Limited benchmark coverage (1 scores).
Confidence: 40% · Freshness: Evaluation date not disclosed · Evidence: 33% · Benchmarks: 33% · Stability: 13%
-
#7 DeepSeek R1 — DeepSeek
Fit score 49 out of 100.
- Strong reasoning (76/100).
- Strong cost efficiency (75/100).
- Open-weight (MIT).
- Caution: Overall fit moderate (49/100) — consider alternatives.
Confidence: 40% · Freshness: Evaluation date not disclosed · Evidence: 67% · Benchmarks: 67% · Stability: 38%
-
#8 Llama 3.1 8B — Meta
Fit score 48 out of 100.
- Strong cost efficiency (98/100).
- Open-weight (Llama 3.1 Community).
- Caution: Overall fit moderate (48/100) — consider alternatives.
- Caution: Limited benchmark coverage (1 scores).
Confidence: 40% · Freshness: Evaluation date not disclosed · Evidence: 33% · Benchmarks: 33% · Stability: 13%
-
#9 Mistral Nemo — Mistral AI
Fit score 47 out of 100.
- Strong cost efficiency (96/100).
- Open-weight (Apache-2.0).
- Caution: Overall fit moderate (47/100) — consider alternatives.
- Caution: Limited benchmark coverage (1 scores).
Confidence: 40% · Freshness: Evaluation date not disclosed · Evidence: 33% · Benchmarks: 33% · Stability: 13%
-
#10 Llama 4 Scout — Meta
Fit score 36 out of 100.
- Strong cost efficiency (94/100).
- Strong ocr (89/100).
- Strong multimodal (76/100).
- Open-weight (Llama 4 Community).
- Caution: Weak coding (33/100).
- Caution: Overall fit moderate (36/100) — consider alternatives.
Confidence: 40% · Freshness: Evaluation date not disclosed · Evidence: 80% · Benchmarks: 80% · Stability: 100%
Limits of this ranking
Coverage is uneven. A model ranks only where ModelRefs holds benchmark-eligible evidence for the capabilities this use case requires, so a strong model with thin evidence can rank low or be absent entirely. Confidence, evidence, and benchmark-coverage figures beside each candidate say how well supported its position is — read them before acting on the order.
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Best Open Source AI Models.