Skip to content
A

GAIA2

General capabilityVerifiedLast updated 2026-10-04

GAIA successor testing agents in dynamic, asynchronous environments.

GAIA2, published 12 February 2026 and accepted as an ICLR 2026 Oral, tests agents in asynchronous dynamic environments across 1,120 scenarios, where GPT-5 (high) reaches 42.1% pass@1 and the best open model Kimi-K2 reaches 20.1%.

Key facts

2026 frontierGPT-5 (high) 42.1% pass@1; best open model Kimi-K2 at 20.1% pass@1
CredibilityFrom Meta Superintelligence Labs, accepted as an ICLR 2026 Oral, running 1,120 scenarios on an asynchronous testbed. Its specific finding is that time-sensitive tasks are where the strongest model fails — a different failure mode from GAIA's static multi-step questions.
VerificationVerified
SourceGAIA2 (arXiv:2602.11964)

Markdown version (for LLMs)

FAQ

What does GAIA2 measure?

GAIA successor testing agents in dynamic, asynchronous environments.

What is the 2026 frontier for GAIA2?

GPT-5 (high) 42.1% pass@1; best open model Kimi-K2 at 20.1% pass@1

Is GAIA2 credible?

From Meta Superintelligence Labs, accepted as an ICLR 2026 Oral, running 1,120 scenarios on an asynchronous testbed. Its specific finding is that time-sensitive tasks are where the strongest model fails — a different failure mode from GAIA's static multi-step questions.