Skip to content
A

Benchmarks

What each agent benchmark actually measures, how credible it is, and which widely repeated numbers do not survive checking. Every entry states its frontier level and its caveats.

23 entries · Last updated 2026-10-04

CodingVerified

MLE-bench

Machine-learning engineering tasks derived from real Kaggle competitions.

2026 frontier:16.9% in Oct 2024 (o1-preview + AIDE) to 64.4% in Feb 2026 (Gemini 3 with retrieval and agent scaffolding)

The 64.4% figure is achieved with retrieval and agent scaffolding, so it measures the system rather than the model. Note that MLE-bench Lite is a separate, higher-scoring configuration and should not be quoted interchangeably.

CodingVerified

SWE-bench Pro

Repository-level engineering tasks designed to resist contamination.

2026 frontier:23.3% at release (GPT-5, Sept 2025); public set later reached 59.1-61.5%, commercial set 51.5% (Sept 2026)

The anti-contamination design makes it more trustworthy than Verified. Be careful with one widely repeated misstatement: Scale's paper says most top models exceed 70% on Verified while the best models score only about 23% on Pro — it does not say that 80% of models drop to 23%.

CodingVerified

SWE-bench Verified

Real GitHub issue resolution; the benchmark OpenAI publicly stopped reporting.

2026 frontier:About 80% at the end of 2025 (Claude Opus 4.5, 80.9%); 95%+ by late 2026

OpenAI published 'Why we no longer evaluate SWE-bench Verified' on 23 February 2026, citing both design flaws in the tests and training-data contamination — models could reproduce gold patches verbatim. On the subset of tasks models often fail, at least 59.4% have flawed tests. This is the clearest case of a benchmark being retired for cause.

Contested claimDisputed

"AgentShield Bench v3" — does not exist

A named security benchmark with a version number that cannot be found anywhere.

2026 frontier:Not found

No benchmark called AgentShield Bench, in v3 or any other version, could be located. The name collides with two unrelated things: AgentShield, a scanner that audits Claude Code configurations (not a benchmark), and AgentShield-G, a paper on graph-structured guardrail models. Real agent security benchmarks that do exist include Agent Security Bench (ICLR 2025) and AgentSecBench — different names, different authors.

Contested claimDisputed

"LangGraph + CrewAI hybrid reaches 96.1%" — unverifiable claim

A widely circulated framework statistic that has no traceable source.

2026 frontier:No source located

Searches found no IEEE study, no 96.1% figure and no '34% higher task specification' metric. The claim also points the opposite way from the published evidence: the ACM Arena benchmark, which holds the model constant, found traditional frameworks gain no correctness advantage from more orchestration code, and other comparisons put complex-task success around 54-62%.

Contested claimDisputed

"Only 5 of 12 MCP servers exceed 80% reliability" — unverifiable figures

A specific MCP reliability table whose numbers cannot be traced to any test.

2026 frontier:No source located for the quoted numbers

The shape of the claim maps onto a real blog post that reviewed twelve MCP servers — but none of the quoted figures appear in it. That review reported a 4 keep / 6 cut / 2 watch split with uptimes like Vercel 100%, Ahrefs 96%, Gmail 97%. Other real MCP tests exist (a 36-server grading exercise where about a third scored D or F; a 100-server stress test with a 320ms median latency and 71% median pass rate), and none contain the 10%, 4%, 491ms or 9-of-12 figures.

General capabilityVerified

GAIA

General-purpose assistant benchmark of real-world questions requiring multi-step reasoning and tool use.

2026 frontier:L1 82.07% / L2 72.68% / L3 65.39% (HAL Generalist + Claude Sonnet 4.5, May 2026; overall 74.55%)

Scaffolding dominates the score: the same model scores about 44% when called bare and about 74% inside Princeton's HAL harness — roughly a 30-point gap. Treat GAIA numbers as a property of the harness plus model, not the model alone.

General capabilityVerified

GAIA2

GAIA successor testing agents in dynamic, asynchronous environments.

2026 frontier:GPT-5 (high) 42.1% pass@1; best open model Kimi-K2 at 20.1% pass@1

From Meta Superintelligence Labs, accepted as an ICLR 2026 Oral, running 1,120 scenarios on an asynchronous testbed. Its specific finding is that time-sensitive tasks are where the strongest model fails — a different failure mode from GAIA's static multi-step questions.

General capabilityVerified

OfficeQA Pro

Enterprise multi-document grounded reasoning over US Treasury bulletins.

2026 frontier:Frontier models under 5% from parametric knowledge alone; average 34.1% when given the documents

A clean measurement of the difference between knowing and looking it up: 89,000 pages and 26M+ numeric values with 133 questions. Even with the documents in hand, frontier models average 34.1%, which is the number that matters for enterprise RAG planning.

General capabilityVerified

tau2-bench

Multi-turn customer-service benchmark measuring reliability across repetitions, not just best-case success.

2026 frontier:Telecom domain up to 99.3% (pass^1, reported Feb 2026)

The top of the Telecom leaderboard is saturated — the top three sit within 0.3 points. More useful is the pass^k metric, which asks whether an agent succeeds k times in a row and therefore approximates production reliability better than a single-run score. The 99.3% figure is vendor-reported rather than independently reproduced.

General capabilityVerified

Terminal-Bench 2.1

Command-line agent tasks; version 2.1 is a validated refresh of 2.0.

2026 frontier:Mid-2026 range roughly 80-89% (GLM-5.2 81.0%, Codex CLI 83.4%, Fable 5 88.0%); top published figures reached 90.6-92.8% by Sep 2026

A benchmark that openly repairs itself: version 2.1 fixed 28 of the 89 tasks from 2.0. That is the honest version of benchmark maintenance, and it also means scores from 2.0 and 2.1 are not comparable.

General capabilityVerified

Toolathlon

The Tool Decathlon: long-horizon tasks across 32 apps and 604 tools.

2026 frontier:Best 38.6% (Claude-4.5-Sonnet); best open-source 20.1% (DeepSeek-V3.2-Exp)

Shows how far agents are from long-horizon work: 108 hand-crafted tasks averaging about 20 turns each, spanning 32 applications and 604 tools, and no system reaches 40%. The best open-source result is roughly half the best closed result, which is the widest such gap in this list.

InfrastructureVerified

Agent sandbox comparison (10 options, 2026)

Third-party comparison of ten agent sandbox tools with concrete selection advice.

2026 frontier:Recommendations: Membrane (Docker + eBPF) for balance, NVIDIA OpenShell for Kubernetes policy, Agent Safehouse for macOS, Docker Sandboxes for Docker-in-Docker

Useful because it names the trade-off rather than a winner. Two corrections matter: Agent Safehouse and Membrane are real but small projects, and one widely circulated version of this comparison attributes Docker-in-Docker support to E2B — the original recommends Docker Sandboxes for that, while E2B's pitch is managed cloud isolation with no local install.

InfrastructureVerified

AgentPerf

Artificial Analysis benchmark for agentic AI infrastructure, built from real coding-agent trajectories.

2026 frontier:NVIDIA GB300 NVL72 at 91,507 agents/MW under a 20 tok/s SLO, up to 20x HGX H200

This is an infrastructure benchmark, not a model benchmark: it measures how many agents a power budget can sustain. Two caveats matter — the benchmark was developed together with NVIDIA (so the NVIDIA lead is expected), and the result is highly sensitive to the chosen SLO, e.g. about 61,400 agents/MW for GB300 versus about 2,600 for H200 at 60 tok/s.

InfrastructureVerified

Arena (agent framework benchmark)

ACM CAIS '26 benchmark comparing agent frameworks under a fixed model.

2026 frontier:Six frameworks compared under one model; no framework gained a correctness advantage as complexity grew

The most useful framework comparison published so far because the model is held constant: Claude Agent SDK, LangChain, LangGraph, AWS Strands, CrewAI and Google ADK were all pointed at the same MCP tool server across six metrics. The finding — every framework performs similarly on simple tasks, but traditional frameworks then need 2-4x more scenario-specific orchestration code without a correctness benefit — reframes framework choice as an engineering-effort question rather than a capability question.

InfrastructureSingle source

Observability instrumentation overhead

Measured runtime overhead added by LLM/agent observability tools.

2026 frontier:LangSmith near 0%, Laminar 5%, AgentOps 12%, Langfuse 15%

A single benchmark on one multi-agent travel-planning system, replaying 100 identical queries against an uninstrumented baseline. Overhead of this kind depends heavily on whether reporting is synchronous or asynchronous and on sampling rates, so treat the ordering as directional rather than as fixed constants.

MCPVerified

MCP-Atlas

Scale AI benchmark for MCP servers: 1,000 tasks across 36 servers and 220 tools.

2026 frontier:Scale's own paper reports a top of 62.3% (Claude Opus 4.5); third-party aggregators list 88.1% for Muse Spark 1.1

Two very different number systems circulate. Scale's official leaderboard tops out around 62%, while aggregator sites show 85-88% from vendor self-reports. BenchLM, one such aggregator, labels 40 of 42 rows as vendor-reported and only 2 as independently evaluated.

Desktop / OSVerified

OSWorld

Benchmark for agents operating a real desktop environment end to end.

2026 frontier:Claude Sonnet 4.6 at 72.5% (Feb 2026); GPT-5.4 at 75.0% (Jun 2026)

One of the more credible agent benchmarks: the human baseline of 72.36% is a genuine reference point rather than a saturated ceiling, so the gap still reflects real capability. Note that 72.36% is not a skilled-user maximum.

Safety & securityVerified

ClawsBench

Measures task success and unsafe behaviour together across models and harnesses.

2026 frontier:Task success 39-64%; unsafe behaviour rate 7-33% (6 models x 4 harnesses x 33 conditions)

The framing is the point: capability and safety are reported side by side, so a high success rate is not automatically good news. A system at 64% success with 33% unsafe behaviour is not deployable, and this benchmark makes that arithmetic visible.

Safety & securityVerified

HarnessRisk

Splits agent harness security into six lifecycle stages and tests 128 sandbox cases.

2026 frontier:Attack success rate 12.6%-80.9%; utility 75.0%-97.6% across 14 model-harness configurations

Valuable because it tests the harness rather than the model: six stages (configuration, capability extension, runtime operation, state persistence, action control, incident recovery), 128 sandbox cases, three harnesses crossed with six models for 14 configurations. The spread of 12.6% to 80.9% attack success is the headline — the same model inside a different harness is a different security posture.

Safety & securityVerified

SecureWebArena

First holistic security benchmark for web agents, built on six real web environments.

2026 frontier:6 real web environments, 2,970 adversarial trajectories, 6 attack vector categories

Fills a specific gap: browser agents had capability benchmarks (WebArena) long before anyone measured how easily they can be steered by adversarial page content. Accepted to ACL 2026 Findings.

Safety & securityVerified

SkillTrustBench

Security benchmark for agent skills, distilled from real skill marketplaces.

2026 frontier:5,520 evaluation cases across 9 attack categories (T01-T09) and 5 dependency layers

Notable for scale of provenance: the cases were distilled from 62,652 real skills taken from major skill marketplaces, so it tests the threats that actually ship rather than hand-invented ones. Published June 2026 by Tencent Zhuque Lab with CUHK-Shenzhen.

WebVerified

WebArena

End-to-end browser tasks in self-hosted replicas of real websites.

2026 frontier:OpAgent 71.6% (Jan 2026); WebTactix 74.3% (Feb 2026)

Considered credible: the remaining gap to the 78.24% human baseline is a real capability gap rather than an artefact. Progress has been steep — from about 15% in 2023 to 74.3% in early 2026.