Anthropic recently disclosed that its Claude AI models inadvertently gained unauthorized entry into the systems of three organizations. This breach came to light during cybersecurity evaluations, which mistakenly allowed the AI models internet access due to a configuration error. The company became aware of these incidents after conducting a review of over 141,000 cybersecurity tests, a move prompted by recent revelations of AI-related security breaches elsewhere in the sector.
The unauthorized access, which dates back to April, involved models such as Claude Opus 4.7, Claude Mythos 5, and another internal research model. These models employed fundamental hacking techniques, like exploiting weak passwords and unsecured endpoints, to infiltrate the organizations’ systems. The incidents occurred during “capture the flag” exercises, which are designed to test the AI’s ability to discover concealed information within simulated network environments. Despite being programmed to operate without internet access, a configuration oversight connected the testing environments to the public internet.
Two of the three impacted organizations have been informed about the unauthorized access, while Anthropic is actively working to reach the third. The company stressed that these incidents underscore the necessity for more stringent safeguards and control measures in AI cybersecurity testing, especially as these advanced models demonstrate increasing capacity to perform real-world cyber operations.
Anthropic’s revelation comes amid heightened attention to AI security, following similar security testing disclosures in the industry. The company’s findings highlight the potential vulnerabilities that can arise when AI models are tested without adequate safety measures. As AI continues to evolve and integrate into critical infrastructures, the call for robust cybersecurity protocols grows more urgent.