Bottom line: AI pentest tools deliver more vulnerability findings, but according to a survey, up to a quarter of them are regularly waste such as duplicate reports or fabricated CVE numbers, turning triage into the new bottleneck.
A survey of 158 security practitioners shows that while AI-powered pentesting tools uncover significantly more vulnerabilities, the manual verification of those findings eats back into the promised time savings. For CISOs, this means automation shifts effort rather than reducing it.
The vendor Pentest-Tools.com surveyed 158 security practitioners in June 2026 who actively use AI tools for vulnerability analysis, including penetration testers, security engineers, DevSecOps specialists, and MSSP staff. Of 147 participants who use AI to generate findings, 87.8 percent stated that a significant portion of the results must be manually reviewed. For 26.5 percent, this regularly affects more than a quarter of all findings. One participant reported a tool that delivered 300 findings, 250 of which turned out to be waste: duplicate reports, non-exploitable SQL injections, and fabricated CVE numbers. Only 20.3 percent of teams have a functioning workflow to review more than 500 AI-generated suspected cases from a single engagement. 38.6 percent say their team is heavily burdened as a result, while 29.7 percent consider the volume unmanageable.
For CISOs, this shifts the calculus around deploying such tools: the bottleneck is no longer finding vulnerabilities but verifying them. According to the survey, how well a team copes with this depends less on company size than on testing frequency. Teams with infrequent testing cycles are less likely to have established triage processes and are more likely to be caught off guard by the sheer volume of data. The biggest source of frustration, cited in around 30 percent of free-text responses, is false-positive reports, invented exploits, and fabricated findings — well ahead of cost or integration issues. Respondents specifically mentioned convincing-looking but non-reproducible reports as well as fabricated CVE numbers.
This shift is also reflected in tool procurement criteria: 63 percent of respondents name the false-positive rate as the most important criterion, 53 percent cite proof of actual exploitability. Cost ranks only third, at 47 percent. AI is used primarily for scanning (74.1 percent) and report writing (69 percent), and considerably less often in areas where a tester interacts live with a system: only 36.7 percent for vulnerability exploitation, and just 25.3 percent for post-exploitation steps. This can be read as a deliberate risk trade-off, since a hallucinated attack step against a production system is barely correctable, whereas an error in a draft report is.
Respondents frequently cite the tools’ lack of understanding of business logic as their biggest remaining weakness: a tool may reliably find SQL injections but fail to recognize that a discount voucher should only be redeemable once per customer. Such logic flaws produce no technical error message and thus remain a task for human testers. At the same time, AI itself is increasingly becoming a test target: 92.4 percent of respondents already test AI-powered applications or plan to do so within twelve months — though for nearly half, only upon specific request. On the topic of shadow AI, i.e. unauthorized AI use by employees, only a third of teams have firmly anchored corresponding risks in their testing scope.
Source: www.it-daily.net · Published August 5, 2026
Lumi AI News — AI-assisted curation pursuant to Art. 50 EU AI Act. Paraphrasing and classification by Lumi News Pipeline v1.8.3.