# GAIA2

> GAIA2, published 12 February 2026 and accepted as an ICLR 2026 Oral, tests agents in asynchronous dynamic environments across 1,120 scenarios, where GPT-5 (high) reaches 42.1% pass@1 and the best open model Kimi-K2 reaches 20.1%.

GAIA successor testing agents in dynamic, asynchronous environments.

- **Category:** General capability
- **2026 frontier:** GPT-5 (high) 42.1% pass@1; best open model Kimi-K2 at 20.1% pass@1
- **Credibility:** From Meta Superintelligence Labs, accepted as an ICLR 2026 Oral, running 1,120 scenarios on an asynchronous testbed. Its specific finding is that time-sensitive tasks are where the strongest model fails — a different failure mode from GAIA's static multi-step questions.
- **Verification:** Verified
- **Source:** [GAIA2 (arXiv:2602.11964)](https://arxiv.org/abs/2602.11964)
- **Last updated:** 2026-10-04
