ModelRefs / SWE-Bench Multimodal Leaderboard — AI Model Scores
SWE-Bench Multimodal Leaderboard — AI Model Scores
Fix UI bugs that require screenshot understanding alongside code. Current leaders, methodology, and citation sources for SWE-Bench Multimodal.
Overview
Fix UI bugs that require screenshot understanding alongside code.
How it is measured: Resolve-rate on JavaScript repos with attached screenshots.
How this benchmark is scored
| Category | coding |
|---|---|
| Maximum score | 100 % resolved |
| Direction | Higher is better |
Primary source: https://www.swebench.com/multimodal.html
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to SWE-Bench Multimodal Leaderboard — AI Model Scores.