SWE-bench Verified

SWE-bench Verified is a benchmark for evaluating AI agents on resolving real-world GitHub issues, with results tracked over time. As of July 2026, the top scores are Claude Mythos 5 at 95.5% and Claude Fable 5 at 95.0%.