Skip to content
A

GAIA

General capabilityVerifiedLast updated 2026-10-04

General-purpose assistant benchmark of real-world questions requiring multi-step reasoning and tool use.

As of May 2026 the GAIA leaderboard showed L1 82.07%, L2 72.68% and L3 65.39% for a HAL-scaffolded Claude Sonnet 4.5, against a human baseline of 92% — but the same model scores roughly 44% without the harness, a 30-point scaffolding gap.

Key facts

2026 frontierL1 82.07% / L2 72.68% / L3 65.39% (HAL Generalist + Claude Sonnet 4.5, May 2026; overall 74.55%)
Human baseline92% (original paper)
CredibilityScaffolding dominates the score: the same model scores about 44% when called bare and about 74% inside Princeton's HAL harness — roughly a 30-point gap. Treat GAIA numbers as a property of the harness plus model, not the model alone.
VerificationVerified
SourceGAIA leaderboard / HAL analysis

Markdown version (for LLMs)

FAQ

What does GAIA measure?

General-purpose assistant benchmark of real-world questions requiring multi-step reasoning and tool use.

What is the 2026 frontier for GAIA?

L1 82.07% / L2 72.68% / L3 65.39% (HAL Generalist + Claude Sonnet 4.5, May 2026; overall 74.55%)

Is GAIA credible?

Scaffolding dominates the score: the same model scores about 44% when called bare and about 74% inside Princeton's HAL harness — roughly a 30-point gap. Treat GAIA numbers as a property of the harness plus model, not the model alone.