Skip to content
A

MLE-bench

CodingVerifiedLast updated 2026-10-04

Machine-learning engineering tasks derived from real Kaggle competitions.

MLE-bench scores rose from 16.9% in October 2024 with o1-preview plus AIDE to 64.4% in February 2026 with Gemini 3 using retrieval and agent scaffolding — a system-level figure, not a bare-model one.

Key facts

2026 frontier16.9% in Oct 2024 (o1-preview + AIDE) to 64.4% in Feb 2026 (Gemini 3 with retrieval and agent scaffolding)
CredibilityThe 64.4% figure is achieved with retrieval and agent scaffolding, so it measures the system rather than the model. Note that MLE-bench Lite is a separate, higher-scoring configuration and should not be quoted interchangeably.
VerificationVerified
SourceMLE-bench (OpenAI paper) / reporting

Markdown version (for LLMs)

FAQ

What does MLE-bench measure?

Machine-learning engineering tasks derived from real Kaggle competitions.

What is the 2026 frontier for MLE-bench?

16.9% in Oct 2024 (o1-preview + AIDE) to 64.4% in Feb 2026 (Gemini 3 with retrieval and agent scaffolding)

Is MLE-bench credible?

The 64.4% figure is achieved with retrieval and agent scaffolding, so it measures the system rather than the model. Note that MLE-bench Lite is a separate, higher-scoring configuration and should not be quoted interchangeably.