AgentPerf
Artificial Analysis benchmark for agentic AI infrastructure, built from real coding-agent trajectories.
AgentPerf, published by Artificial Analysis with NVIDIA on 12 June 2026, reported 91,507 agents/MW for NVIDIA's GB300 NVL72 at a 20 tok/s SLO — up to 20x the HGX H200 — but the figure drops to about 61,400 versus 2,600 at a 60 tok/s SLO.
Contested claim: Quoting 'up to 20x' without stating the SLO, which changes the ratio by roughly 5x.
Key facts
| 2026 frontier | NVIDIA GB300 NVL72 at 91,507 agents/MW under a 20 tok/s SLO, up to 20x HGX H200 |
|---|---|
| Credibility | This is an infrastructure benchmark, not a model benchmark: it measures how many agents a power budget can sustain. Two caveats matter — the benchmark was developed together with NVIDIA (so the NVIDIA lead is expected), and the result is highly sensitive to the chosen SLO, e.g. about 61,400 agents/MW for GB300 versus about 2,600 for H200 at 60 tok/s. |
| Verification | Verified |
| Source | Artificial Analysis hardware benchmarks |
FAQ
What does AgentPerf measure?
Artificial Analysis benchmark for agentic AI infrastructure, built from real coding-agent trajectories.
What is the 2026 frontier for AgentPerf?
NVIDIA GB300 NVL72 at 91,507 agents/MW under a 20 tok/s SLO, up to 20x HGX H200
Is AgentPerf credible?
This is an infrastructure benchmark, not a model benchmark: it measures how many agents a power budget can sustain. Two caveats matter — the benchmark was developed together with NVIDIA (so the NVIDIA lead is expected), and the result is highly sensitive to the chosen SLO, e.g. about 61,400 agents/MW for GB300 versus about 2,600 for H200 at 60 tok/s.
Is the commonly cited claim about AgentPerf accurate?
Quoting 'up to 20x' without stating the SLO, which changes the ratio by roughly 5x.