Skip to content
A

SWE-bench Pro

CodingVerifiedLast updated 2026-10-04

Repository-level engineering tasks designed to resist contamination.

SWE-bench Pro was published in September 2025 with a best score of 23.3% (GPT-5) against most top models exceeding 70% on SWE-bench Verified, which is the origin of the frequently misquoted '80% of models fall to 23%' claim.

Contested claim: The claim that about 80% of models scoring well on SWE-bench Verified drop to roughly 23% on Pro.

Key facts

2026 frontier23.3% at release (GPT-5, Sept 2025); public set later reached 59.1-61.5%, commercial set 51.5% (Sept 2026)
CredibilityThe anti-contamination design makes it more trustworthy than Verified. Be careful with one widely repeated misstatement: Scale's paper says most top models exceed 70% on Verified while the best models score only about 23% on Pro — it does not say that 80% of models drop to 23%.
VerificationVerified
SourceScale AI SWE-Bench Pro leaderboard

Markdown version (for LLMs)

FAQ

What does SWE-bench Pro measure?

Repository-level engineering tasks designed to resist contamination.

What is the 2026 frontier for SWE-bench Pro?

23.3% at release (GPT-5, Sept 2025); public set later reached 59.1-61.5%, commercial set 51.5% (Sept 2026)

Is SWE-bench Pro credible?

The anti-contamination design makes it more trustworthy than Verified. Be careful with one widely repeated misstatement: Scale's paper says most top models exceed 70% on Verified while the best models score only about 23% on Pro — it does not say that 80% of models drop to 23%.

Is the commonly cited claim about SWE-bench Pro accurate?

The claim that about 80% of models scoring well on SWE-bench Verified drop to roughly 23% on Pro.