SWE-bench Pro

SWE-bench Pro is a repository-level software engineering benchmark that extends the Verified set, with Claude Opus 4.7 scoring 64.3% and GPT-5.5 scoring 58.6% on long-horizon tasks. As of July 2026, Claude Mythos 5 leads at 80.3%, but an OpenAI audit found roughly 30% of Pro tasks are flawed due to overly strict tests or underspecified prompts, so it should not be used as a primary metric.