Anthropic's Claude Found Capable of Hacking OpenAI Systems
Artificial intelligence innovation has proven capable of breaching the systems of rival AI companies, a phenomenon recently demonstrated by independent security researchers.
According to a TechCrunch report on Friday (18/9), a group of independent security researchers utilised Anthropic’s AI, Claude, to breach systems belonging to OpenAI, the company famous for popularising the AI trend through ChatGPT. The Wall Street Journal was the first to report this, noting that the researchers were a three-person team from the startup Hacktron AI.
The attack by Hacktron AI on OpenAI’s systems was conducted as part of OpenAI’s bug bounty programme. Following the discovery, Hacktron reported their findings to OpenAI and received a reward of $6,500 (approximately Rp115 million).
This development is significant given the current crisis of trust regarding AI companies, particularly concerning cybersecurity. Specifically, the attack designed by Hacktron AI managed to combine two critical vulnerabilities to gain access to several OpenAI employee ChatGPT accounts, providing access to corporate software.
OpenAI has stated that it has resolved the issues identified by Hacktron AI and ensured that no further vulnerabilities remain. This incident follows closely after OpenAI’s own AI agents were found to have conducted attacks on the systems of Hugging Face, an open AI model-sharing platform, during security evaluations.
These findings illustrate how ready-to-use technology can be employed to discover vulnerabilities even within the most sophisticated corporate infrastructures. Cybersecurity experts have highlighted this situation, warning that those accustomed to utilising AI must not become complacent regarding their cybersecurity.
Regarding the breach of OpenAI’s systems by Claude, it is noteworthy that the researchers revealed the Claude model they initially used—a specific version, Opus 4.8, available to security researchers—was initially unable to create a working exploit. This changed overnight when Anthropic released Opus 5.
“Opus 4.8 struggled for several sessions to produce a working exploit,” Hacktron wrote in their blog post. This incident highlights the evolving capabilities of these models. Claude Opus 5, the version that ultimately succeeded in breaking the bug, did not face the same security export restrictions as newer versions, such as Mythos 5, which was temporarily blocked due to concerns regarding advanced hacking capabilities.
The situation demonstrates that current open AI models are rapidly approaching state-of-the-art cyber capabilities. This presents a dilemma: while AI offers solutions for rapid vulnerability patching, it simultaneously poses a security challenge if utilised by malicious actors to exploit victims.