ModelRefs / MathVista Leaderboard — AI Model Scores

MathVista Leaderboard — AI Model Scores

Math reasoning over visual contexts (charts, figures, geometry). Current leaders, methodology, and citation sources for MathVista.

Overview

Math reasoning over visual contexts (charts, figures, geometry).

How it is measured: TestMini split; mean accuracy.

What this benchmark measures

  • visual perception
  • mathematical reasoning over visual contexts
  • structured answer generation

Relevant to:

  • visual-math evaluation portfolios
  • chart and figure reasoning screening for finance-document workflows

Failure modes it exercises:

  • incorrect visual interpretation
  • incorrect mathematical reasoning
  • answer-extraction failure

Method and its limits

Accuracy on visual mathematical questions drawn from 28 existing multimodal datasets and three newly created datasets, with result interpretation pinned to the dataset split, prompt, visual inputs, and answer-extraction protocol.

  • The aggregate combines heterogeneous task types, skills, and source datasets, so overall accuracy can hide category-specific weaknesses.
  • OCR, captioning, chain-of-thought, program tools, prompt format, and answer extraction can materially affect results.

Dataset

Dataset
MathVista
Type
visual mathematical reasoning questions across charts, plots, tables, diagrams, document images, puzzles, and scientific figures
Freshness
unknown

The primary sources describe 6,141 examples from 28 existing datasets and three new datasets; this does not establish coverage of current financial documents or accounting tasks.

How to read this score

Use category-level and problem-level results under a pinned protocol; do not infer financial-document, accounting, or production reliability from aggregate accuracy.

Similarity to real tasks: Medium — The benchmark includes charts, plots, tables, diagrams, document images, and scientific figures, but does not reproduce an organization's financial records, controls, or review workflow.

Data contamination risk: Unknown — The benchmark consolidates public datasets and ModelRefs has no model-run-specific training-overlap assessment.

Benchmark gaming risk: Unknown — Prompting, auxiliary OCR or captions, tools, answer extraction, and benchmark-targeted tuning can affect results.

What you still need to test yourself

  • Test the organization's actual invoices, reports, charts, tables, layouts, currencies, tax formats, and review rules.
  • Measure field and figure accuracy, source traceability, abstention, reviewer correction, latency, cost, and control failures separately.

This benchmark supports decisions about:

  • Screen multimodal systems for visual mathematical reasoning under matched prompts, inputs, tools, and answer extraction.
  • Identify visual, OCR, chart, table, and reasoning failure categories that should appear in internal finance-document tests.

Limitations

  • MathVista is not an accounting, compliance, audit, fraud, forecasting, or financial-control benchmark.
  • Visual mathematical accuracy does not prove reliable extraction, source traceability, policy adherence, or human-review outcomes on target documents.

Sources reviewed 2026-07-02. Dataset freshness remains unknown; pin the dataset split, repository revision, prompt, auxiliary inputs, tools, and answer-extraction protocol.

Sources

How this benchmark is scored

Categorymultimodal
Maximum score100 % accuracy
DirectionHigher is better

Primary source: https://mathvista.github.io/

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to MathVista Leaderboard — AI Model Scores.