# OSWorld

> Agents crossed the OSWorld human baseline of 72.36% in February 2026 when Claude Sonnet 4.6 reached 72.5%, and GPT-5.4 reached 75.0% by June 2026 — one of the few agent benchmarks where the human reference line has been overtaken.

Benchmark for agents operating a real desktop environment end to end.

- **Category:** Desktop / OS
- **2026 frontier:** Claude Sonnet 4.6 at 72.5% (Feb 2026); GPT-5.4 at 75.0% (Jun 2026)
- **Human baseline:** 72.36% (original paper)
- **Credibility:** One of the more credible agent benchmarks: the human baseline of 72.36% is a genuine reference point rather than a saturated ceiling, so the gap still reflects real capability. Note that 72.36% is not a skilled-user maximum.
- **Verification:** Verified
- **Source:** [OSWorld](http://osworld-v1.xlang.ai/)
- **Last updated:** 2026-10-04
