Skip to content

Evaluation of AI Pentesting Agents under Realistic Conditions

The Bottom Line: A structured evaluation protocol with multiple complex test environments and LLM-powered vulnerability detection enables more realistic assessment of AI pentesting agents beyond classical benchmark scenarios.

Researchers have introduced a new evaluation protocol for AI-based pentesting systems that measures their performance in complex, real-world attack scenarios rather than in simplified benchmark environments. The protocol enables, for the first time, an operationally relevant comparison of different AI pentesting agents.

Previous evaluations of AI pentesting agents have focused on defined target categories such as capture-the-flag, remote code execution, or exploit reproduction. While these simplified or narrowly defined test scenarios demonstrate limited technical capabilities, they do not capture the complexity of real penetration tests: open-ended exploration, multiple attack surfaces, diverse vulnerability classes, and strategic decision-making under uncertainty.

The new evaluation method shifts the focus from pure task completion to validated vulnerability detection. The protocol integrates structured ground-truth data (expert-annotated reference information), LLM-based semantic matching procedures for vulnerability identification, bipartite matching for assessment under realistic ambiguities, continuous maintenance of ground truth, and repeated and cumulative evaluation of stochastic agents. Additionally, efficiency metrics and optimized test suites are provided to enable sustainable experiments.

For CISOs, the protocol provides a better foundation for assessing AI pentesting solutions before production deployment. Rather than relying on marketing promises, security teams can now make decisions based on realistic complexity regarding which agent is suitable for their specific infrastructure. This reduces the risk of deploying agentic systems whose actual performance under real conditions is unknown.

The researchers have made the evaluation protocol, reference ground-truth data, and code available at https://github.com/ethiack/ethibench to enable reproducibility and further research.


Source: arxiv.org · Published 13 July 2026
Lumi AI News — AI-assisted curation according to Art. 50 EU AI Act. Paraphrase and classification by Lumi News Pipeline v1.7.3.

Share on: