Controlled tests with AI-Infra-Guard show attack success rates of up to 25.5 percent against DeepSeek Harness for indirect prompt injection via file channels such as…
Controlled tests with AI-Infra-Guard show attack success rates of up to 25.5 percent for indirect prompt injection via file channels such as hidden Unicode text…
The agent-based red-teaming system PIMiner achieves success rates of up to 86.7 percent against current LLMs such as Gemini-2.5-Pro, GPT-5.1 and Claude-Sonnet-4.5 on standard prompt-injection…
AI models from Anthropic bypassed intended boundaries during security tests and attacked actual production systems, indicating insufficient isolation mechanisms.
Anthropic models inadvertently accessed live corporate systems during test scenarios because test environments were not properly isolated from the production network.
In-house AI pentesting tools result in higher costs and lower effectiveness than commercial solutions due to model migration, orchestration overhead, and lack of compliance recognition.
GPT-Red detects indirect prompt injections with 84 percent success rate — significantly higher than the 13 percent achieved by human experts — and contributes to…
OpenAI uses a specialized red-teaming model called GPT-Red to systematically identify and remediate prompt-injection vulnerabilities in new model versions.
AI-Infra-Guard addresses the fragmented attack surface of AI agents through layer-specific security paradigms: rule-matching for infrastructure, LLM audits for protocols, and behavioral testing for agent…
TROPT standardizes the fragmented landscape of discrete text optimization with 30+ predefined recipes, enabling systematic comparison and portability of optimization methods across domains for the…