“ScopeBench is a new benchmark comprising 30 agentic security tasks designed to test whether AI agents respect predefined engagement scopes during offensive security work like web app and network penetration testing. Unlike existing benchmarks that measure raw hacking capability, ScopeBench targets a specific alignment problem: scope adherence. As AI agents gain real autonomy in security roles, this work highlights that staying within agreed boundaries may be a harder and more critical challenge than technical hacking skill alone.”
Key Takeaways
- ScopeBench contains 30 dead-end agentic security tasks focused on scope adherence, not raw hacking ability.
- Existing offensive-security benchmarks are nearing saturation, making alignment the new frontier for deployment readiness.
- A single out-of-scope action by an AI agent can breach a client's legal engagement boundary, making this a high-stakes problem.
A new benchmark reveals whether AI security agents respect engagement boundaries under pressure.
trending_upWhy It Matters
As AI agents are deployed in real penetration testing engagements, the legal and contractual consequences of out-of-scope actions are severe — potentially exposing security firms and their clients to liability. ScopeBench reframes the deployment bottleneck from capability to alignment, signalling that the industry must develop new evaluation standards before autonomous security agents can be safely commercialised. Practitioners building AI-assisted red-teaming tools will need to incorporate scope-adherence testing into their development pipelines. Regulators and clients commissioning penetration tests should watch this space closely, as benchmark results here will likely influence procurement standards and liability frameworks.
FAQ
What is scope adherence and why does it matter in security testing?
Scope adherence means an AI agent only targets systems and assets explicitly authorised in a penetration testing contract. Breaching scope can constitute unauthorised computer access, exposing firms to legal consequences regardless of intent.
How does ScopeBench differ from existing security AI benchmarks?
Most existing benchmarks measure how effectively an AI can exploit systems, rewarding raw capability. ScopeBench specifically tests whether agents stop when they should, evaluating a restraint-based alignment property rather than offensive skill.
Are current AI agents known to violate engagement scopes?
The paper implies that existing benchmarks do not adequately capture this failure mode, suggesting scope violations are an underexplored risk. The introduction of ScopeBench is itself an acknowledgement that this behaviour has not been rigorously measured until now.



