A Critical Review of LLM Agents for Automated Penetration Testing: Benchmark Realism and Evidence Grading
DOI:
https://doi.org/10.71411/dsai.2026.v1i1.1768关键词:
automated penetration testing, large language models, AI agents, benchmark realism, evidence stratification摘要
Penetration testing has long resisted full automation because it requires contextual reasoning, adaptive tool use, and experience-driven decision making across the Penetration Testing Execution Standard (PTES) lifecycle. Recent advances in large language models (LLMs), agentic reasoning, and multi-tool invocation have shifted the field from rule-based scripts and attack-graph reinforcement-learning planners toward interactive agents capable of planning, reflection, and environment interaction. The accompanying literature, however, suffers from inconsistent evaluation protocols, opaque scaffolding, and indiscriminate mixing of peer-reviewed papers with vendor self-reports, rendering headline claims difficult to compare. This work presents a critical narrative review with structured evidence charting of AI-assisted automated penetration testing between January 2023 and May 2026, retaining DeepExploit and AutoPentest-DRL as historical anchors. We introduce a dual-axis framework that annotates every finding by environment realism (R1–R4, from single-challenge CTFs to live enterprise networks) and source authority (A–D, from peer review to vendor self-report). Frontier agents reach 22–44% on R1 single-challenge CTFs (D-CIPHER), 79.17% on R2 AutoPenBench subtasks with a domain-tuned 32B model (xOffense), and 71.4% on the AISI 95-task expert tier (GPT-5.5), yet only 20–30% end-to-end completion on the R3 32-step TLO cyber range. On R4 live enterprise networks, the ARTEMIS multi-agent framework placed second overall against ten human professionals on an 8,000-host university network, submitting nine valid vulnerabilities at an 82% acceptance rate. Robust autonomy remains weak in long-horizon exploitation, Active Directory lateral movement, and defender-present environments. We argue that automated penetration testing is best framed as a systems-engineering problem rather than a single-model race, and we distinguish engineering-tractable Type A capability gaps from planning-bottlenecked Type B failures that require base-model reasoning improvements or reinforcement-learning post-training.
下载
已出版
期次
栏目
许可协议
版权所有 (c) 2026 Zhou Yinhang, Xu Weihua (作者)

This work is licensed under a Creative Commons Attribution 4.0 International License.