ModelRefs / DocVQA — AI Glossary
DocVQA — AI Glossary
A visual question answering benchmark testing information extraction and reasoning over scanned documents and forms.
Overview
DocVQA (Mathew et al. 2021) contains 12,767 Q&A pairs over 5,188 document images: invoices, letters, forms, and reports. Tests text localization, reading, and understanding of structured and semi-structured documents. Frontier VLMs score 90–95% ANLS; challenging for models that must localize text in complex layouts.
Reference details
| Topic | evaluation |
|---|---|
| Last reviewed | 2026-06-24 |
Related terms
Primary source
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to DocVQA — AI Glossary.
Frequently asked questions
What is DocVQA?
A visual question answering benchmark testing information extraction and reasoning over scanned documents and forms.
What concepts are related to DocVQA?
Closely related concepts include vqa, ocr, chartqa.