Skip to content
A

HarnessRisk

Safety & securityVerifiedLast updated 2026-10-04

Splits agent harness security into six lifecycle stages and tests 128 sandbox cases.

HarnessRisk, published 18 August 2026, evaluates agent harness security across six lifecycle stages and 128 sandbox cases in 14 model-harness configurations, finding attack success rates from 12.6% to 80.9%.

Key facts

2026 frontierAttack success rate 12.6%-80.9%; utility 75.0%-97.6% across 14 model-harness configurations
CredibilityValuable because it tests the harness rather than the model: six stages (configuration, capability extension, runtime operation, state persistence, action control, incident recovery), 128 sandbox cases, three harnesses crossed with six models for 14 configurations. The spread of 12.6% to 80.9% attack success is the headline — the same model inside a different harness is a different security posture.
VerificationVerified
SourceHarnessRisk (arXiv:2608.17597)

Markdown version (for LLMs)

FAQ

What does HarnessRisk measure?

Splits agent harness security into six lifecycle stages and tests 128 sandbox cases.

What is the 2026 frontier for HarnessRisk?

Attack success rate 12.6%-80.9%; utility 75.0%-97.6% across 14 model-harness configurations

Is HarnessRisk credible?

Valuable because it tests the harness rather than the model: six stages (configuration, capability extension, runtime operation, state persistence, action control, incident recovery), 128 sandbox cases, three harnesses crossed with six models for 14 configurations. The spread of 12.6% to 80.9% attack success is the headline — the same model inside a different harness is a different security posture.