ModelRefs / Best Cheap AI Models
Best Cheap AI Models
Most cost-efficient AI models for production workloads — ranked by quality per dollar.
Overview
This page answers: What is the cheapest model that still works in production?
Candidates are ranked against the startup-mvp use case using current ModelRefs evidence. The ordering below was computed when this page was built and is recomputed on every deploy; with JavaScript enabled it is re-ranked against the live catalogue on load.
How this ranking is produced
Recommendation engine phase-2.0.0, 16 benchmarks evaluated, aggregate freshness fresh. Computed from the canonical ModelRefs registry when this page was built on 2026-09-04, and recomputed on every deploy.
Each candidate below is a provisional fit signal under the stated constraints, not a guarantee, a certification, or a final ranking. Fit scores compare models against this use case's capability weights using current ModelRefs evidence; they are not accuracy rates, benchmark results, or production-readiness claims.
What we measure for Startup MVPs
- Cost Efficiency
- 90%
- Latency
- 70%
- Instruction Following
- 70%
Weights derived from the ModelRefs capability ontology. Scores sourced from primary benchmark leaderboards where available; expansion entries carry lower confidence (0.65).
Ranked candidates
-
#1 Mixtral 8x7B — Mistral AI
Fit score 61 out of 100.
- Strong cost efficiency (91/100).
- Strong reasoning (83/100).
- Strong instruction following (83/100).
- Pricing $0.45/$0.45 per Mtok.
- Caution: Limited benchmark coverage (1 scores).
Capability evidence
- Cost Efficiency — evidence confidence 40% — Top-tier cost efficiency (91/100).
- Latency — evidence confidence 40% — Underperforms on latency (0/100).
- Instruction Following — evidence confidence 40% — Top-tier instruction following (83/100). via mt-bench
Confidence: 40% · Freshness: Evaluation date not disclosed · Evidence: 33% · Benchmarks: 33% · Stability: 13%
-
#2 Llama 3.1 8B — Meta
Fit score 60 out of 100.
- Strong cost efficiency (98/100).
- Pricing $0.05/$0.08 per Mtok.
- Caution: Limited benchmark coverage (1 scores).
Capability evidence
- Cost Efficiency — evidence confidence 40% — Top-tier cost efficiency (98/100).
- Latency — evidence confidence 40% — Underperforms on latency (0/100).
Confidence: 40% · Freshness: Evaluation date not disclosed · Evidence: 0% · Benchmarks: 0% · Stability: 13%
-
#3 Llama 3.1 70B — Meta
Fit score 59 out of 100.
- Strong cost efficiency (86/100).
- Strong reasoning (84/100).
- Strong instruction following (84/100).
- Pricing $0.59/$0.79 per Mtok.
- Caution: Limited benchmark coverage (1 scores).
Capability evidence
- Cost Efficiency — evidence confidence 40% — Top-tier cost efficiency (86/100).
- Latency — evidence confidence 40% — Underperforms on latency (0/100).
- Instruction Following — evidence confidence 40% — Top-tier instruction following (84/100).
Confidence: 40% · Freshness: Evaluation date not disclosed · Evidence: 0% · Benchmarks: 0% · Stability: 13%
-
#4 Mistral Nemo — Mistral AI
Fit score 58 out of 100.
- Strong cost efficiency (96/100).
- Pricing $0.15/$0.15 per Mtok.
- Caution: Limited benchmark coverage (1 scores).
Capability evidence
- Cost Efficiency — evidence confidence 40% — Top-tier cost efficiency (96/100).
- Latency — evidence confidence 40% — Underperforms on latency (0/100).
Confidence: 40% · Freshness: Evaluation date not disclosed · Evidence: 0% · Benchmarks: 0% · Stability: 13%
-
#5 Phi-4 — Microsoft
Fit score 57 out of 100.
- Strong cost efficiency (100/100).
- Strong multilingual (81/100).
- Pricing $0/$0 per Mtok.
- Caution: Weak hallucination resistance (3/100).
Capability evidence
- Cost Efficiency — evidence confidence 40% — Top-tier cost efficiency (100/100).
- Latency — evidence confidence 40% — Underperforms on latency (0/100).
Confidence: 40% · Freshness: Stale 0% · Evidence: 0% · Benchmarks: 0% · Stability: 88%
-
#6 Command R — Cohere
Fit score 56 out of 100.
- Strong cost efficiency (91/100).
- Pricing $0.15/$0.6 per Mtok.
- Caution: Limited benchmark coverage (1 scores).
Capability evidence
- Cost Efficiency — evidence confidence 40% — Top-tier cost efficiency (91/100).
- Latency — evidence confidence 40% — Underperforms on latency (0/100).
Confidence: 40% · Freshness: Evaluation date not disclosed · Evidence: 0% · Benchmarks: 0% · Stability: 13%
-
#7 BGE-M3 — BAAI
Fit score 39 out of 100.
- Strong cost efficiency (100/100).
- Strong multilingual (95/100).
- Strong rag suitability (95/100).
- Pricing $0/$0 per Mtok.
- Caution: Overall fit moderate (39/100) — consider alternatives.
Capability evidence
- Cost Efficiency — evidence confidence 40% — Top-tier cost efficiency (100/100).
- Latency — evidence confidence 40% — Underperforms on latency (0/100).
- Instruction Following — evidence confidence 40% — Underperforms on instruction following (0/100).
Confidence: 40% · Freshness: Evaluation date not disclosed · Evidence: 0% · Benchmarks: 0% · Stability: 38%
-
#8 Qwen2.5-VL 72B — Alibaba
Fit score 39 out of 100.
- Strong cost efficiency (100/100).
- Strong ocr (88/100).
- Strong multimodal (78/100).
- Pricing $0/$0 per Mtok.
- Caution: Overall fit moderate (39/100) — consider alternatives.
Capability evidence
- Cost Efficiency — evidence confidence 40% — Top-tier cost efficiency (100/100).
- Latency — evidence confidence 40% — Underperforms on latency (0/100).
- Instruction Following — evidence confidence 40% — Underperforms on instruction following (0/100).
Confidence: 40% · Freshness: Evaluation date not disclosed · Evidence: 0% · Benchmarks: 0% · Stability: 50%
-
#9 InternVL2.5 78B — OpenGVLab
Fit score 39 out of 100.
- Strong cost efficiency (100/100).
- Pricing $0/$0 per Mtok.
- Caution: Overall fit moderate (39/100) — consider alternatives.
- Caution: Limited benchmark coverage (1 scores).
Capability evidence
- Cost Efficiency — evidence confidence 40% — Top-tier cost efficiency (100/100).
- Latency — evidence confidence 40% — Underperforms on latency (0/100).
- Instruction Following — evidence confidence 40% — Underperforms on instruction following (0/100).
Confidence: 40% · Freshness: Evaluation date not disclosed · Evidence: 0% · Benchmarks: 0% · Stability: 13%
-
#10 Text Embedding 3 Small — OpenAI
Fit score 39 out of 100.
- Strong cost efficiency (100/100).
- Strong rag suitability (77/100).
- Pricing $0.02/$0 per Mtok.
- Caution: Overall fit moderate (39/100) — consider alternatives.
- Caution: Limited benchmark coverage (2 scores).
Capability evidence
- Cost Efficiency — evidence confidence 40% — Top-tier cost efficiency (100/100).
- Latency — evidence confidence 40% — Underperforms on latency (0/100).
- Instruction Following — evidence confidence 40% — Underperforms on instruction following (0/100).
Confidence: 40% · Freshness: Evaluation date not disclosed · Evidence: 0% · Benchmarks: 0% · Stability: 25%
Knowledge Graph signal
Independent cross-validation from the ModelRefs semantic graph. Models below were identified via graph traversal of benchmark to capability to use-case edges — a separate signal from the recommendation engine above.
Capability fit describes how strongly a model's measured capabilities match this use case. It is not an accuracy rate or production-readiness guarantee. Evidence confidence describes how complete and well-supported the evidence behind that fit is; missing or stack-level requirements lower confidence.
- #1 BGE-M3 — evidence confidence 65%
- #2 Qwen2.5-VL 72B — evidence confidence 65%
- #3 InternVL2.5 78B — evidence confidence 65%
- #4 Text Embedding 3 Small — evidence confidence 65%
- #5 Text Embedding Ada 002 — evidence confidence 65%
- #6 Text Embedding 3 Large — evidence confidence 65%
Limits of this ranking
Coverage is uneven. A model ranks only where ModelRefs holds benchmark-eligible evidence for the capabilities this use case requires, so a strong model with thin evidence can rank low or be absent entirely. Confidence, evidence, and benchmark-coverage figures beside each candidate say how well supported its position is — read them before acting on the order.
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Best Cheap AI Models.