# GAIA

> As of May 2026 the GAIA leaderboard showed L1 82.07%, L2 72.68% and L3 65.39% for a HAL-scaffolded Claude Sonnet 4.5, against a human baseline of 92% — but the same model scores roughly 44% without the harness, a 30-point scaffolding gap.

General-purpose assistant benchmark of real-world questions requiring multi-step reasoning and tool use.

- **Category:** General capability
- **2026 frontier:** L1 82.07% / L2 72.68% / L3 65.39% (HAL Generalist + Claude Sonnet 4.5, May 2026; overall 74.55%)
- **Human baseline:** 92% (original paper)
- **Credibility:** Scaffolding dominates the score: the same model scores about 44% when called bare and about 74% inside Princeton's HAL harness — roughly a 30-point gap. Treat GAIA numbers as a property of the harness plus model, not the model alone.
- **Verification:** Verified
- **Source:** [GAIA leaderboard / HAL analysis](https://awesomeagents.ai/leaderboards/gaia-benchmark-leaderboard/)
- **Last updated:** 2026-10-04
