Skip to content

OpenAI Develops GPT-Red for Automated Red-Teaming Against Prompt Injections

Bottom line: GPT-Red detects indirect prompt injections with 84 percent success rate — significantly higher than the 13 percent achieved by human experts — and contributes to developing safer model generations.

OpenAI has released a specialized AI model called GPT-Red that automatically identifies security vulnerabilities through prompt injections and attacks on instruction hierarchy. The system deliberately simulates new attack variants to make AI models more robust before their release.

OpenAI deploys GPT-Red in a controlled training loop as an automated attack model. The system initially identifies vulnerabilities in models under review. Once these successfully block existing attack patterns, GPT-Red deliberately generates new, more complex variants of prompt injections and hierarchy attacks. The attack data generated in this way is fed directly into the training and security validation of more robust model generations.

In internal tests on indirect prompt injections, GPT-Red achieved an 84 percent success rate in uncovering security vulnerabilities. By comparison, human security experts identified only 13 percent of vulnerabilities in the same scenarios. The attack patterns generated by GPT-Red were used, among other things, in the development of GPT-5.6 Sol. For direct prompt injections, GPT-5.6 Sol showed six times fewer successful manipulations in the most demanding internal test procedures compared to the strongest production model four months earlier.

The need for such security tools grows as autonomous AI agents are integrated into production environments. These systems can independently access web browsers, applications, emails, and local files — the potential attack surface for manipulated commands is thereby significantly enlarged. OpenAI views GPT-Red as a scaling tool for increasing the scope and difficulty level of security testing, and employs it as a complement to manual reviews by external experts, multilayered security barriers, monitoring systems, and vulnerability disclosure programs.


Source: www.it-daily.net · Published July 17, 2026
Lumi AI News — AI-assisted curation pursuant to Art. 50 EU AI Act. Paraphrase and classification by Lumi News Pipeline v1.7.3.

Share on: