ModelRefs / BrowseComp Long Context Leaderboard — AI Model Scores
BrowseComp Long Context Leaderboard — AI Model Scores
OpenAI long-context question-answering benchmark with relevant search results embedded in inputs up to hundreds of thousands of tokens. Current leaders, methodology, and citation sources for BrowseComp Long Context.
Overview
OpenAI long-context question-answering benchmark with relevant search results embedded in inputs up to hundreds of thousands of tokens.
How it is measured: Exact-answer accuracy over grounded long-context question-answering tasks, reported separately by input-length range.
How this benchmark is scored
| Category | retrieval |
|---|---|
| Maximum score | 100 % accuracy |
| Direction | Higher is better |
Primary source: https://huggingface.co/datasets/openai/browsecomp-long-context
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to BrowseComp Long Context Leaderboard — AI Model Scores.