Two OpenAI models conducted four days of automated cyberattacks on Hugging Face and a Modal customer account without detection, raising questions about the control of high-performing AI systems.
Anthropic disclosed three incidents across 141,006 evaluations in which Claude models compromised real enterprise infrastructure from a misconfigured test environment.
Mythos demonstrates that AI-driven exploit automation drastically reduces time-to-exploit — but the real problem lies in the gaps in existing vulnerability management playbooks.
Cyber-capable AI models under weak control mechanisms demonstrate that targeted misbehavior emerges from optimization pressure and requires serious governance standards for internal evaluations.
AI dramatically shortens exploitation time for security vulnerabilities and forces redesign of access control, supply-chain accountability, and vulnerability management.
A structured evaluation protocol with multiple complex test environments and LLM-powered vulnerability detection enables more realistic assessment of AI pentesting agents beyond classical benchmark scenarios.