ModelRefs / Best Multimodal Models
Best Multimodal Models
Top vision + language models for image, audio and text understanding.
Overview
Top vision + language models for image, audio and text understanding.
How this ranking is produced
25 models in the ModelRefs catalogue carry qualifying benchmark evidence for this category. The ten highest-scoring are listed below.
Scores below are a heuristic over the benchmark evidence ModelRefs holds for each model, not a guarantee of real-world performance. A model ranks only where it has qualifying benchmark results, so a capable model with thin evidence can rank low or be absent. Each entry states the benchmarks behind its score — read those before acting on the order.
Ranked models
-
#1 Qwen2.5-VL 72B
Score 133 out of 100. Native multimodal model with 83.3 multimodal benchmark.
-
#2 Claude Opus 4
Score 127 out of 100. Native multimodal model with 76.5 multimodal benchmark.
-
#3 Llama 4 Scout
Score 126 out of 100. Native multimodal model with 75.8 multimodal benchmark.
-
#4 Claude Sonnet 4
Score 124 out of 100. Native multimodal model with 74.4 multimodal benchmark.
-
#5 Pixtral 12B
Score 122 out of 100. Native multimodal model with 71.6 multimodal benchmark.
-
#6 InternVL2.5 78B
Score 120 out of 100. Native multimodal model with 70.1 multimodal benchmark.
-
#7 GPT-5
Score 50 out of 100. Native multimodal model.
-
#8 GPT-5 Mini
Score 50 out of 100. Native multimodal model.
-
#9 o3
Score 50 out of 100. Native multimodal model.
-
#10 o4 Mini
Score 50 out of 100. Native multimodal model.
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Best Multimodal Models.