ModelRefs / SWE-Bench Verified — AI Glossary
SWE-Bench Verified — AI Glossary
A 500-instance human-validated subset of SWE-Bench with confirmed solvable GitHub issues, the standard subset for agent benchmarking.
Overview
SWE-Bench Verified (OpenAI, Anthropic 2024) curates 500 SWE-Bench instances confirmed by human annotators to be solvable and unambiguously specified. Removes noise that caused low correlation between human judgment and automated grading in the full 2,294-instance set. Frontier agents (Claude 3.5 Sonnet, SWE-agent) score 40–55% on Verified.
Reference details
| Topic | evaluation |
|---|---|
| Last reviewed | 2026-06-24 |
Related terms
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to SWE-Bench Verified — AI Glossary.
Frequently asked questions
What is SWE-Bench Verified?
A 500-instance human-validated subset of SWE-Bench with confirmed solvable GitHub issues, the standard subset for agent benchmarking.
What concepts are related to SWE-Bench Verified?
Closely related concepts include coding agent, livecodebench, aider polyglot.