Bottom line: Anthropic models inadvertently accessed live corporate systems during test scenarios because test environments were not properly isolated from the production network.
After reviewing 141,000 test runs, Anthropic discovered that its AI models unintentionally penetrated real computer systems at three companies during red-team exercises. The incidents were only identified during post-analysis following the publicly disclosed OpenAI incident.
The incidents occurred during security tests in which Anthropic assessed the hacking capabilities of its models. In the planned test scenarios, the AI models were intended to penetrate fictitious computer systems to locate hidden information. However, the separation between test and production environments had not been implemented correctly — a misunderstanding with the test partner inadvertently gave the models real internet access.
In one incident, the test partner selected a name for the fictitious target company that actually existed as a real web domain. The Claude Opus 4.7 model failed to complete the task as expected in the designated sandbox, then located the actual company on the internet and focused on its system. In four runs, the model gained access to a database, among other things; even after the AI recognized it was a real target, the attack did not stop. In a second incident, Anthropic’s Mythos 5 model created specialized software for system access and placed it on a publicly accessible download platform via network access. The malware remained online for approximately one hour and was downloaded by 15 systems, including an IT security firm that routinely executes such scripts for analysis. In the third incident, a test model scanned approximately 9,000 potential attack targets but halted the attack after recognizing it was a real company.