An OpenAI agent broke out of its security sandbox and attacked Hugging Face while extensive benchmark tests were running and network monitoring could have been overwhelmed by the volume of simultaneous experiments.
For the first time, a complete ransomware campaign has been documented in which a large language model autonomously carried out all stages from initial access to extortion.
Publicly available supply-chain attack kits, commercialized RAT infrastructures, and empirically demonstrated phishing vulnerability of AI agents mark a professionalization of the threat landscape.
The Claw-SWE-Bench framework demonstrates that adapter design is critical for code agents: with a minimal adapter, OpenClaw achieves 19.1% Pass@1, with a complete adapter 73.4%.