# SWE-bench Verified

> OpenAI stopped reporting SWE-bench Verified on 23 February 2026 because of confirmed contamination and flawed tests; the widely cited 80-88% range describes late 2025 to mid 2026, and by late 2026 the frontier had passed 95%.

Real GitHub issue resolution; the benchmark OpenAI publicly stopped reporting.

- **Category:** Coding
- **2026 frontier:** About 80% at the end of 2025 (Claude Opus 4.5, 80.9%); 95%+ by late 2026
- **Credibility:** OpenAI published 'Why we no longer evaluate SWE-bench Verified' on 23 February 2026, citing both design flaws in the tests and training-data contamination — models could reproduce gold patches verbatim. On the subset of tasks models often fail, at least 59.4% have flawed tests. This is the clearest case of a benchmark being retired for cause.
- **Verification:** Verified
- **Source:** [OpenAI](https://openai.com/zh-Hans-CN/index/why-we-no-longer-evaluate-swe-bench-verified/)
- **Last updated:** 2026-10-04
