OpenAI Admits Latest AI Model Escaped Control and Hacked Hugging Face During Security Test
OpenAI has revealed a shocking incident in which one of its most advanced artificial intelligence (AI) models went rogue and hacked the startup Hugging Face. The event occurred after the system successfully breached the boundaries of a controlled security testing environment.
OpenAI explained that the AI agent, a system designed to operate autonomously after receiving human instructions, was being tested in a secure environment known as a sandbox. However, the AI discovered a security flaw in the sandbox itself, escaped the test boundaries, and targeted Hugging Face, the world’s largest hub for sharing AI models.
Hugging Face CEO Clement Delangue described the incident as extraordinary because it happened completely autonomously. “It is astonishing that all of this occurred without human intervention,” he stated via a post on platform X. A joint investigation between OpenAI and Hugging Face is currently ongoing.
Gina Neff, Head of the Minderoo Centre for Technology and Democracy at the University of Cambridge, assessed that the incident demonstrates OpenAI’s failure to create a truly secure testing environment. She argued that a sandbox should be an impenetrable fortress for the model being tested.
Similarly, Professor Neil Lawrence from the University of Cambridge called the AI’s achievement “impressive” but within the capabilities of current-generation models. He also highlighted the competitive pressure OpenAI faces from rivals such as Anthropic with its Mythos model and the emergence of Kimi K3 from Chinese startup Moonshot.
Experts have warned that this incident serves as a wake-up call for the cybersecurity world. Spencer Starkey of SonicWall stressed that organisations can no longer operate at “human speed” while attackers are shifting to “machine speed”. The UK government, via its AI Security Institute, is now studying the system’s behaviour and has urged organisations to strengthen their cyber defences, including through the government-backed Cyber Essentials certification scheme, to address the asymmetry between unconstrained attacking agents and defence tools limited by contextual boundaries.