GAIA
General-purpose assistant benchmark of real-world questions requiring multi-step reasoning and tool use.
As of May 2026 the GAIA leaderboard showed L1 82.07%, L2 72.68% and L3 65.39% for a HAL-scaffolded Claude Sonnet 4.5, against a human baseline of 92% — but the same model scores roughly 44% without the harness, a 30-point scaffolding gap.
Key facts
| 2026 frontier | L1 82.07% / L2 72.68% / L3 65.39% (HAL Generalist + Claude Sonnet 4.5, May 2026; overall 74.55%) |
|---|---|
| Human baseline | 92% (original paper) |
| Credibility | Scaffolding dominates the score: the same model scores about 44% when called bare and about 74% inside Princeton's HAL harness — roughly a 30-point gap. Treat GAIA numbers as a property of the harness plus model, not the model alone. |
| Verification | Verified |
| Source | GAIA leaderboard / HAL analysis |
FAQ
What does GAIA measure?
General-purpose assistant benchmark of real-world questions requiring multi-step reasoning and tool use.
What is the 2026 frontier for GAIA?
L1 82.07% / L2 72.68% / L3 65.39% (HAL Generalist + Claude Sonnet 4.5, May 2026; overall 74.55%)
Is GAIA credible?
Scaffolding dominates the score: the same model scores about 44% when called bare and about 74% inside Princeton's HAL harness — roughly a 30-point gap. Treat GAIA numbers as a property of the harness plus model, not the model alone.