ModelRefs / DocVQA Leaderboard — AI Model Scores

DocVQA Leaderboard — AI Model Scores

Document visual question answering on scanned business docs. Current leaders, methodology, and citation sources for DocVQA.

Overview

Document visual question answering on scanned business docs.

How it is measured: ANLS over test set.

What this benchmark measures

  • document-image question answering
  • OCR-dependent document understanding
  • layout- and structure-sensitive answer extraction

Relevant to:

  • document question-answering screening
  • document-intelligence evaluation portfolios
  • visual extraction workflow evaluation
  • invoice, vendor-document, contract, and audit-evidence failure-case screening
  • completion-readiness test design for invoice, AP, EHR, legal, and evidence-synthesis workflows

Failure modes it exercises:

  • incorrect answers from document images
  • failure to use document layout and structure
  • OCR and answer-normalization errors

Method and its limits

Answer natural-language questions about document images using the pinned DocVQA task, split, and evaluation metric; report whether OCR, layout annotations, retrieval, or other auxiliary inputs are available.

  • The original benchmark focuses on single-page document images and question answering rather than full document parsing or schema extraction.
  • ANLS and accuracy are sensitive to answer normalization and do not expose every factual, layout, or provenance error.
  • OCR system, image resolution, auxiliary text, task edition, and document-domain mix can materially change results.

Dataset

Dataset
DocVQA
Type
scanned document images with natural-language questions and short answers
Freshness
aging

The primary paper reports 50,000 questions over more than 12,000 document images. Later DocVQA challenges and multi-page variants are separate scopes and must not be mixed with the original task.

How to read this score

Read the score with task edition, split, metric, OCR condition, and document categories. Use document- and question-type errors rather than one aggregate score for workflow decisions.

Similarity to real tasks: Medium — The dataset uses scanned document images and human questions, but the original single-page task does not reproduce long document collections, custom schemas, handwriting, access controls, or downstream automation.

Data contamination risk: Unknown — The public dataset and challenge materials may be present in training corpora; no model-specific overlap audit is registered.

Benchmark gaming risk: Unknown — OCR choice, document-specific tuning, auxiliary text, answer normalization, and task-edition selection can affect results.

What you still need to test yourself

  • Test representative document families, layouts, scan qualities, languages, tables, handwriting, multi-page context, and required output schemas.
  • Measure field-level accuracy, source grounding, missing-field handling, confidence, human handoff, privacy controls, latency, and cost.
  • Test actual invoice fields, vendor evidence, amendments, retention metadata, source reliability, validation rules, and reviewer corrections.
  • Define deployment-specific acceptance, monitoring, escalation, rollback, and qualified-review gates before treating document performance as workflow evidence.

This benchmark supports decisions about:

  • Screen multimodal systems for bounded question answering over document images.
  • Identify layout, OCR, document-type, and answer-normalization failures that require workload-specific extraction tests.
  • Inform invoice, vendor-intake, contract-review, and evidence-collection test cohorts without treating QA accuracy as workflow validation.
  • Seed document cohorts and failure taxonomies for workflow acceptance testing without satisfying completion-readiness by itself.

Limitations

  • DocVQA does not establish reliable schema extraction, multi-page reasoning, handwriting support, or downstream business-process safety.
  • A question-answering score does not guarantee complete field recall, correct provenance, calibrated confidence, or privacy compliance.
  • It does not establish invoice validity, vendor safety, contract meaning, audit sufficiency, control effectiveness, or permission to act.
  • It cannot by itself make an invoice, AP, EHR, legal, or clinical workflow a completion candidate.

Sources reviewed 2026-07-02. Dataset freshness remains unknown; challenge news and site updates do not establish an original-corpus freshness cutoff.

Sources

How this benchmark is scored

Categoryvision
Maximum score100 ANLS
DirectionHigher is better

Primary source: https://www.docvqa.org/

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to DocVQA Leaderboard — AI Model Scores.