Skip to content

PIMiner: Agent-Based Red Teaming Reveals High Success Rates for Prompt Injection

Bottom line: The agent-based red-teaming system PIMiner achieves success rates of up to 86.7 percent against current LLMs such as Gemini-2.5-Pro, GPT-5.1 and Claude-Sonnet-4.5 on standard prompt-injection benchmarks, without any retraining.

Researchers present PIMiner, an agent-based system for automated red-teaming of prompt-injection attacks that can be transferred to new target LLMs without retraining. In tests, it achieves success rates of up to 86.7 percent against current models such as Gemini-2.5-Pro, GPT-5.1 and Claude-Sonnet-4.5.

PIMiner is an agent-based system for red-teaming prompt-injection attacks against LLM agents. Unlike existing state-of-the-art methods that rely on reinforcement learning and thereby produce attacker models that transfer poorly to new target models, PIMiner builds a strategy library from scratch during training, using a sequence of dataset–target-model pairs. At test time, this library can be transferred directly to a previously unseen target LLM without requiring additional training. Per test case, the system needs only a few requests to the target agent — according to the authors, around ten.

The measured success rates (Attack Success Rate, ASR) are relevant for security leaders because they show that current, near-production models are not sufficiently hardened against prompt-injection attacks. On the IPIArena benchmark, PIMiner achieves an ASR of 76.2 percent against Gemini-2.5-Pro, 61.9 percent against GPT-5.1 and 42.9 percent against Claude-Sonnet-4.5. On AgentDojo, the figures are 86.7 percent against Gemini-2.5-Pro, 53.3 percent against GPT-5.1 and 40.0 percent against Claude-Sonnet-4.5. These figures also highlight differences in the robustness of individual models against injection attacks, with Claude-Sonnet-4.5 showing the lowest attacker success rates on both benchmarks.

For CISOs who deploy LLM agents in production environments or are planning to do so, the work provides a concrete tool for their own risk assessment: a transferable strategy library makes it possible to test agent deployments against prompt injection with comparatively little effort before they go live. At the same time, the generated attack data provides training material to improve defense mechanisms. Anyone operating LLM agents with access to external tools, files or APIs should consider the described methodology as a starting point for their own red-teaming processes, since the authors highlight the low number of requests required per test case as practical for repeated use against changing target models.


Source: arxiv.org · Published August 4, 2026
Lumi AI News — AI-assisted curation pursuant to Art. 50 EU AI Act. Paraphrasing and classification by Lumi News Pipeline v1.8.3.

Share on: