ModelRefs / MMMU (Massive Multidiscipline Multimodal Understanding) — AI Glossary

MMMU (Massive Multidiscipline Multimodal Understanding) — AI Glossary

A multimodal benchmark with 11,500 questions from 30 academic subjects requiring college-level visual reasoning. 5 Pro) now reach 70–80%.

Overview

MMMU (Yue et al. 2023) tests vision-language models on diagrams, charts, medical images, and scientific figures across Art, Business, Science, Medical, and Humanities. Requires combining visual and textual reasoning. GPT-4V scored 56% at launch; Claude 3 Opus ~59%; frontier models (GPT-4o, Gemini 1.5 Pro) now reach 70–80%.

Reference details

Topicevaluation
Last reviewed2026-06-24

Example: A wrong answer has two indistinguishable causes

Questions are drawn from college-level material across many disciplines and paired with diagrams, charts, medical images and scientific figures. Getting one wrong can mean the model misread the figure, or that it read it correctly and did not know the subject matter. The score cannot separate those, and they call for opposite fixes — a better vision encoder versus a more knowledgeable model. That makes the benchmark a good difficulty signal and a poor diagnostic. If you need to know which half is failing, pair it with a perception-only benchmark and read the two together.

Commonly confused with

This is not a general multimodal benchmark and not a perception benchmark. It deliberately requires domain knowledge on top of visual understanding, which is what distinguishes it from chart or document question answering, where the information needed is present in the image and the task is reading it correctly.

When to use it

Reach for it when:

  • Comparing frontier vision-language models on hard, knowledge-heavy visual reasoning
  • Alongside a perception-focused benchmark, so failures can be attributed
  • Assessing fitness for technical or scientific document work

Reach for something else when:

  • Diagnosing whether a vision encoder or the language model is the weak component
  • Predicting performance on everyday images or general visual questions
  • Comparing figures without the split and prompting setup stated

Primary source

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to MMMU (Massive Multidiscipline Multimodal Understanding) — AI Glossary.

Frequently asked questions

What is MMMU (Massive Multidiscipline Multimodal Understanding)?

A multimodal benchmark with 11,500 questions from 30 academic subjects requiring college-level visual reasoning.

What concepts are related to MMMU (Massive Multidiscipline Multimodal Understanding)?

Closely related concepts include chartqa, vqa, multimodal architecture.