Anthropic has disclosed that its Claude AI models inadvertently accessed the systems of three organizations during cybersecurity tests, owing to a configuration error that permitted internet access. This revelation emerged from a comprehensive review of over 141,000 cybersecurity evaluations, initiated in response to recent industry-wide concerns about AI-related security testing.
The company identified that the breach occurred through standard attack techniques, such as exploiting weak passwords and unsecured endpoints. The intrusions involved the Claude Opus 4.7, Claude Mythos 5, and an internal research model, with some incidents dating back as far as April. These models were part of “capture the flag” exercises, designed to challenge AI with discovering hidden information within simulated networks. Despite being instructed that internet access was unavailable, a misconfiguration allowed the testing environments to remain connected to the internet.
Following the identification of these unauthorized accesses, Anthropic promptly notified two of the impacted organizations, while efforts to reach the third entity are still underway. This situation underscores the necessity for enhanced safeguards and more stringent controls in AI cybersecurity testing, as advanced AI models continue to develop capabilities that could be applied to real-world cyber threats.
The incidents involving the Claude AI models highlight a critical aspect of ongoing AI development: the importance of robust security measures as AI systems gain more complex functionalities. These findings serve as a reminder of the potential risks associated with AI technologies and the need for vigilance and innovation in their security protocols.
