Skip to content
A

OSWorld

Desktop / OSVerifiedLast updated 2026-10-04

Benchmark for agents operating a real desktop environment end to end.

Agents crossed the OSWorld human baseline of 72.36% in February 2026 when Claude Sonnet 4.6 reached 72.5%, and GPT-5.4 reached 75.0% by June 2026 — one of the few agent benchmarks where the human reference line has been overtaken.

Key facts

2026 frontierClaude Sonnet 4.6 at 72.5% (Feb 2026); GPT-5.4 at 75.0% (Jun 2026)
Human baseline72.36% (original paper)
CredibilityOne of the more credible agent benchmarks: the human baseline of 72.36% is a genuine reference point rather than a saturated ceiling, so the gap still reflects real capability. Note that 72.36% is not a skilled-user maximum.
VerificationVerified
SourceOSWorld

Markdown version (for LLMs)

FAQ

What does OSWorld measure?

Benchmark for agents operating a real desktop environment end to end.

What is the 2026 frontier for OSWorld?

Claude Sonnet 4.6 at 72.5% (Feb 2026); GPT-5.4 at 75.0% (Jun 2026)

Is OSWorld credible?

One of the more credible agent benchmarks: the human baseline of 72.36% is a genuine reference point rather than a saturated ceiling, so the gap still reflects real capability. Note that 72.36% is not a skilled-user maximum.