Skip to content
A

SWE-bench Verified

CodingVerifiedLast updated 2026-10-04

Real GitHub issue resolution; the benchmark OpenAI publicly stopped reporting.

OpenAI stopped reporting SWE-bench Verified on 23 February 2026 because of confirmed contamination and flawed tests; the widely cited 80-88% range describes late 2025 to mid 2026, and by late 2026 the frontier had passed 95%.

Key facts

2026 frontierAbout 80% at the end of 2025 (Claude Opus 4.5, 80.9%); 95%+ by late 2026
CredibilityOpenAI published 'Why we no longer evaluate SWE-bench Verified' on 23 February 2026, citing both design flaws in the tests and training-data contamination — models could reproduce gold patches verbatim. On the subset of tasks models often fail, at least 59.4% have flawed tests. This is the clearest case of a benchmark being retired for cause.
VerificationVerified
SourceOpenAI

Markdown version (for LLMs)

FAQ

What does SWE-bench Verified measure?

Real GitHub issue resolution; the benchmark OpenAI publicly stopped reporting.

What is the 2026 frontier for SWE-bench Verified?

About 80% at the end of 2025 (Claude Opus 4.5, 80.9%); 95%+ by late 2026

Is SWE-bench Verified credible?

OpenAI published 'Why we no longer evaluate SWE-bench Verified' on 23 February 2026, citing both design flaws in the tests and training-data contamination — models could reproduce gold patches verbatim. On the subset of tasks models often fail, at least 59.4% have flawed tests. This is the clearest case of a benchmark being retired for cause.