# Arena (agent framework benchmark)

> An ACM CAIS '26 benchmark published 26 May 2026 fixed Claude Sonnet 4.5 across six agent frameworks pointed at one MCP tool server and found that as task complexity grows, traditional frameworks require 2-4x more scenario-specific orchestration code without gaining any correctness advantage.

ACM CAIS '26 benchmark comparing agent frameworks under a fixed model.

- **Category:** Infrastructure
- **2026 frontier:** Six frameworks compared under one model; no framework gained a correctness advantage as complexity grew
- **Credibility:** The most useful framework comparison published so far because the model is held constant: Claude Agent SDK, LangChain, LangGraph, AWS Strands, CrewAI and Google ADK were all pointed at the same MCP tool server across six metrics. The finding — every framework performs similarly on simple tasks, but traditional frameworks then need 2-4x more scenario-specific orchestration code without a correctness benefit — reframes framework choice as an engineering-effort question rather than a capability question.
- **Verification:** Verified
- **Source:** [ACM CAIS '26 (DOI 10.1145/3786335.3813233)](https://dlnext.acm.org/doi/10.1145/3786335.3813233)
- **Last updated:** 2026-10-04
