ModelRefs / SWE-Bench Verified Leaderboard — AI Model Scores
SWE-Bench Verified Leaderboard — AI Model Scores
Resolve real GitHub issues end-to-end with passing test suite. Current leaders, methodology, and citation sources for SWE-Bench Verified.
Overview
Resolve real GitHub issues end-to-end with passing test suite.
How it is measured: Human-verified subset; resolved-rate on full repository context.
What this benchmark measures
- repository-level software issue resolution
- patch validation
Relevant to:
- coding-agent evaluation
- repository maintenance evaluation
- evaluation-pipeline and observability regression-case design
Failure modes it exercises:
- incorrect patches
- test regressions
- failure to resolve a documented issue
Method and its limits
Resolved-rate on a human-verified subset of real repository issues using repository context and test validation.
- Results depend on the agent scaffold, tools, execution environment, and inference budget.
- Passing benchmark tests does not establish security, maintainability, or review quality.
Dataset
- Dataset
- SWE-bench Verified
- Type
- real GitHub issues with repository test suites
- Freshness
- aging
The record identifies the 500-instance SWE-bench Verified release. Repository, language, and issue-type coverage should be checked at the source; the release date is not a guarantee that issue content postdates model training.
How to read this score
Resolved-rate is useful when harness and budget are held constant; cross-system comparisons require matched conditions.
Similarity to real tasks: High — Tasks use real repository issues and tests, while agent harness, environment, and tool configuration still affect transfer.
Data contamination risk: High — The benchmark is based on public GitHub issues and repositories. Human verification improves task quality, but ModelRefs cannot rule out training overlap for any model run without source-scoped decontamination evidence.
Benchmark gaming risk: High — Scores depend on scaffold, tools, execution environment, cost or step budgets, and benchmark-specific agent tuning; cross-run comparisons require matched conditions.
What you still need to test yourself
- Run representative issues from the target languages, repositories, dependencies, and security constraints.
- Review patch maintainability, security, scope, and human acceptance beyond test passage.
- Verify trace completeness, version attribution, tool and environment failures, regression sensitivity, incident linkage, human adjudication, and rollback behavior in the actual evaluation stack.
This benchmark supports decisions about:
- Compare repository-level issue resolution under matched harness, tool, budget, and environment conditions.
- Assess whether a coding system can coordinate changes across an existing codebase and pass benchmark tests.
- Seed versioned coding-agent regression cohorts for an evaluation or observability workflow while keeping internal release criteria separate.
Limitations
- Does not cover every language, repository type, security risk, or software-development workflow.
- A resolved issue is not equivalent to production-ready code.
- A benchmark delta does not prove that an evaluation pipeline or observability system detects, explains, or safely responds to production regressions.
- Every score must identify the exact model, Verified subset or compatible subset, scaffold, tools, inference budget, and source; model training overlap remains unknown unless source-scoped evidence says otherwise.
Sources re-reviewed 2026-07-11. SWE-bench Verified is treated as an aging 2024 static subset; the 2024-08-13 release is provenance, not an issue-corpus freshness cutoff.
Sources
- SWE-bench SWE-bench · accessed 2026-07-11
- SWE-bench: Can Language Models Resolve Real-world GitHub Issues? ICLR 2024 / SWE-bench authors · accessed 2026-07-11
How this benchmark is scored
| Category | coding |
|---|---|
| Maximum score | 100 % resolved |
| Direction | Higher is better |
Primary source: https://www.swebench.com/
Published results
| Model | Score | Run date | Source |
|---|---|---|---|
| GPT-5 | 74.2 | 2026-05-01 | Aggregated public reports |
| Claude Opus 4 | 72.5 | 2026-05-01 | Aggregated public reports |
| DeepSeek R1 | 49.2 | 2026-05-01 | Aggregated public reports |
| GPT-5 Mini | 49 | 2026-05-01 | Aggregated public reports |
Each score reflects the protocol and date of its own source run. Results from different harnesses are not directly comparable.
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to SWE-Bench Verified Leaderboard — AI Model Scores.