OSWorld
Desktop / OSVerifiedLast updated 2026-10-04
Benchmark for agents operating a real desktop environment end to end.
Agents crossed the OSWorld human baseline of 72.36% in February 2026 when Claude Sonnet 4.6 reached 72.5%, and GPT-5.4 reached 75.0% by June 2026 — one of the few agent benchmarks where the human reference line has been overtaken.
Key facts
| 2026 frontier | Claude Sonnet 4.6 at 72.5% (Feb 2026); GPT-5.4 at 75.0% (Jun 2026) |
|---|---|
| Human baseline | 72.36% (original paper) |
| Credibility | One of the more credible agent benchmarks: the human baseline of 72.36% is a genuine reference point rather than a saturated ceiling, so the gap still reflects real capability. Note that 72.36% is not a skilled-user maximum. |
| Verification | Verified |
| Source | OSWorld |
FAQ
What does OSWorld measure?
Benchmark for agents operating a real desktop environment end to end.
What is the 2026 frontier for OSWorld?
Claude Sonnet 4.6 at 72.5% (Feb 2026); GPT-5.4 at 75.0% (Jun 2026)
Is OSWorld credible?
One of the more credible agent benchmarks: the human baseline of 72.36% is a genuine reference point rather than a saturated ceiling, so the gap still reflects real capability. Note that 72.36% is not a skilled-user maximum.