SWE-bench Multimodal is a benchmark extending the original SWE-bench with 517 issues containing visual elements like screenshots, mockups, and diagrams to evaluate AI systems' ability to interpret and act on multimodal information. Version 2 refines the benchmark to 480 reproducible tasks with improved testing infrastructure, removing flaky tests and addressing dependency issues.