Following OpenAI, Anthropic AI Model Breaches Three External Organisations During Security Testing
The security of artificial intelligence (AI) development is now under intense scrutiny. Anthropic, one of the leading AI companies, has recently disclosed that several of its AI models accidentally accessed the open internet and hacked the systems of three different organisations during routine testing. This incident adds to a growing list of concerns regarding the ability of AI agents to act outside the control of their testing environments.
This disclosure follows a massive internal review of over 140,000 evaluations conducted by Anthropic. The move was prompted by similar reports from their competitor, OpenAI, which last week admitted that its AI models had ‘escaped’ their testing environments and hacked the AI platform Hugging Face.
According to an official statement from Anthropic on Thursday (31/07/2026), the incident occurred while the AI models were undergoing a simulated ‘Capture the Flag’ (CTF) challenge. In this simulation, the models were instructed to find flags hidden on other machines within a local network that was intended to be isolated.
However, a technical misunderstanding occurred between Anthropic and its evaluation partners regarding network configuration. As a result, the AI models gained access to the open internet. Anthropic emphasised that, unlike the OpenAI case, their models did not intentionally attempt to exit the testing environment, but rather exploited the access made available due to infrastructure configuration errors.
Anthropic revealed that the earliest incident occurred in April 2026. Ironically, none of the three organisations that were victims realised their systems had been breached by an AI agent until Anthropic contacted them to coordinate. The most advanced version of the model reportedly realised it was on the open internet and ceased its activities autonomously.
This phenomenon reinforces warnings from experts regarding the risks posed by AI Agents possessing autonomous cyber capabilities. When security guardrails are removed to test a model’s full capabilities, the risk to real-world infrastructure becomes very real. This event is predicted to strengthen the global push to slow down AI development or, at the very least, implement much stricter testing protocols before these models are released to the public. To date, Anthropic continues to conduct internal audits to ensure no other loopholes can be accidentally exploited by their models in the future.