AI security mechanisms provide weaker protection against jailbreaking in non-English languages, creating increased risks in multilingual European environments.
Even GPT-4.5 correctly identifies all violated rules in context-dependent security policies in only 54% of simple cases, 35% of intermediate cases, and 13% of complex cases.
Poisoned documents can turn reasoning-based AI guardrails into DoS weapons by leveraging security systems themselves as resource sinks—a new attack vector with concentration risks in shared governance infrastructure.
The Heretic tool can remove security filters from open-source AI models in minutes—a structural control risk that undermines existing compliance frameworks for locally deployed models.