# ClawsBench

> ClawsBench, first posted 6 April 2026, tested 6 models across 4 agent harnesses and 33 conditions and found task success rates of 39-64% alongside unsafe behaviour rates of 7-33%.

Measures task success and unsafe behaviour together across models and harnesses.

- **Category:** Safety & security
- **2026 frontier:** Task success 39-64%; unsafe behaviour rate 7-33% (6 models x 4 harnesses x 33 conditions)
- **Credibility:** The framing is the point: capability and safety are reported side by side, so a high success rate is not automatically good news. A system at 64% success with 33% unsafe behaviour is not deployable, and this benchmark makes that arithmetic visible.
- **Verification:** Verified
- **Source:** [ClawsBench (arXiv:2604.05172)](https://arxiv.org/abs/2604.05172)
- **Last updated:** 2026-10-04
