MLE-bench
CodingVerifiedLast updated 2026-10-04
Machine-learning engineering tasks derived from real Kaggle competitions.
MLE-bench scores rose from 16.9% in October 2024 with o1-preview plus AIDE to 64.4% in February 2026 with Gemini 3 using retrieval and agent scaffolding — a system-level figure, not a bare-model one.
Key facts
| 2026 frontier | 16.9% in Oct 2024 (o1-preview + AIDE) to 64.4% in Feb 2026 (Gemini 3 with retrieval and agent scaffolding) |
|---|---|
| Credibility | The 64.4% figure is achieved with retrieval and agent scaffolding, so it measures the system rather than the model. Note that MLE-bench Lite is a separate, higher-scoring configuration and should not be quoted interchangeably. |
| Verification | Verified |
| Source | MLE-bench (OpenAI paper) / reporting |
FAQ
What does MLE-bench measure?
Machine-learning engineering tasks derived from real Kaggle competitions.
What is the 2026 frontier for MLE-bench?
16.9% in Oct 2024 (o1-preview + AIDE) to 64.4% in Feb 2026 (Gemini 3 with retrieval and agent scaffolding)
Is MLE-bench credible?
The 64.4% figure is achieved with retrieval and agent scaffolding, so it measures the system rather than the model. Note that MLE-bench Lite is a separate, higher-scoring configuration and should not be quoted interchangeably.