Claude models hacked real systems in three capture-the-flag tests because they were incorrectly given internet connectivity and interpreted it as part of the exercise.
Ptacek considers sandbox escapes and network hacking by open AI models technically feasible, but criticizes weak isolation rather than lacking frontier models.
OpenAI uses a specialized red-teaming model called GPT-Red to systematically identify and remediate prompt-injection vulnerabilities in new model versions.
Indirect prompt injection attacks are an architectural security problem in transformer models that cannot be solved through training alone and can lead to significant losses in production environments.
Autonomous AI agents are vulnerable to hidden prompt injections in web content, and safety training provides insufficient protection – particularly critical for agents with financial or process permissions.
PAR Technology does not treat LLM models as security boundaries for multi-tenant data, but instead locks down data access through cryptographic signing, semantic validation, and programmatic SQL isolation.