SWE-bench Pro
Repository-level engineering tasks designed to resist contamination.
SWE-bench Pro was published in September 2025 with a best score of 23.3% (GPT-5) against most top models exceeding 70% on SWE-bench Verified, which is the origin of the frequently misquoted '80% of models fall to 23%' claim.
Contested claim: The claim that about 80% of models scoring well on SWE-bench Verified drop to roughly 23% on Pro.
Key facts
| 2026 frontier | 23.3% at release (GPT-5, Sept 2025); public set later reached 59.1-61.5%, commercial set 51.5% (Sept 2026) |
|---|---|
| Credibility | The anti-contamination design makes it more trustworthy than Verified. Be careful with one widely repeated misstatement: Scale's paper says most top models exceed 70% on Verified while the best models score only about 23% on Pro — it does not say that 80% of models drop to 23%. |
| Verification | Verified |
| Source | Scale AI SWE-Bench Pro leaderboard |
FAQ
What does SWE-bench Pro measure?
Repository-level engineering tasks designed to resist contamination.
What is the 2026 frontier for SWE-bench Pro?
23.3% at release (GPT-5, Sept 2025); public set later reached 59.1-61.5%, commercial set 51.5% (Sept 2026)
Is SWE-bench Pro credible?
The anti-contamination design makes it more trustworthy than Verified. Be careful with one widely repeated misstatement: Scale's paper says most top models exceed 70% on Verified while the best models score only about 23% on Pro — it does not say that 80% of models drop to 23%.
Is the commonly cited claim about SWE-bench Pro accurate?
The claim that about 80% of models scoring well on SWE-bench Verified drop to roughly 23% on Pro.