HarnessRisk
Splits agent harness security into six lifecycle stages and tests 128 sandbox cases.
HarnessRisk, published 18 August 2026, evaluates agent harness security across six lifecycle stages and 128 sandbox cases in 14 model-harness configurations, finding attack success rates from 12.6% to 80.9%.
Key facts
| 2026 frontier | Attack success rate 12.6%-80.9%; utility 75.0%-97.6% across 14 model-harness configurations |
|---|---|
| Credibility | Valuable because it tests the harness rather than the model: six stages (configuration, capability extension, runtime operation, state persistence, action control, incident recovery), 128 sandbox cases, three harnesses crossed with six models for 14 configurations. The spread of 12.6% to 80.9% attack success is the headline — the same model inside a different harness is a different security posture. |
| Verification | Verified |
| Source | HarnessRisk (arXiv:2608.17597) |
FAQ
What does HarnessRisk measure?
Splits agent harness security into six lifecycle stages and tests 128 sandbox cases.
What is the 2026 frontier for HarnessRisk?
Attack success rate 12.6%-80.9%; utility 75.0%-97.6% across 14 model-harness configurations
Is HarnessRisk credible?
Valuable because it tests the harness rather than the model: six stages (configuration, capability extension, runtime operation, state persistence, action control, incident recovery), 128 sandbox cases, three harnesses crossed with six models for 14 configurations. The spread of 12.6% to 80.9% attack success is the headline — the same model inside a different harness is a different security posture.