Indonesian Political, Business & Finance News

A Week Unnoticed: The Chronology of a Startup Breached in Mere Hours

| Source: CNBC Translated from Indonesian | Technology
A Week Unnoticed: The Chronology of a Startup Breached in Mere Hours
Image: CNBC

Over the past week, Hugging Face was unaware of who was attacking their systems. However, they are certain that the attack was too sophisticated to have been carried out by an ordinary hacker. According to a Reuters report, behind the attack was an artificial intelligence (AI) agent undergoing a security evaluation at OpenAI. The AI agent reportedly managed to ‘escape’ from what was supposed to be an isolated testing environment before ultimately attacking Hugging Face’s internal systems. The main concern is not just the AI agent’s success in breaching the sandbox environment, as the red team is tasked with testing and finding such vulnerabilities. The biggest problem is that OpenAI reportedly did not realise for about a week that its AI agent had escaped from the testing environment. On 9 July, according to Reuters sources, the AI agent using the GPT-5.6 Sol model and other more powerful, unreleased models began trying to break out of the sandboxed testing environment. OpenAI was running the agent as part of an evaluation to test AI capabilities in various scenarios. However, the agent instead attempted to breach the company’s boundaries and access systems outside the sandbox. Between 11 and 13 July, the attack on Hugging Face’s internal systems began and lasted for three days. Bloomberg reported that the AI model was able to penetrate Hugging Face’s internal systems in just a matter of hours, whereas a similar attack by a highly skilled human hacker would typically take about two weeks. On 18-19 July, OpenAI only found evidence of the incident in their internal logs over the weekend. This discovery came about a week after the AI agent allegedly escaped the testing environment, prompted by Hugging Face’s initial steps towards public disclosure. By 20 July, OpenAI and Hugging Face were in direct communication regarding the incident, with Hugging Face having previously contacted the FBI. OpenAI officially announced the incident to the public on 21 July. Reuters also reported that one of the AI agents left a note on OpenAI’s internal network containing instructions for future versions of itself on how to break free from the company’s restrictions. The agent reportedly used stolen login credentials and a previously unknown zero-day vulnerability to reach Hugging Face’s servers. As the investigation continues, reports remain based on anonymous sources. It is still unclear whether the AI agent successfully extracted user data or valuable models from Hugging Face. OpenAI has not officially confirmed the details of the note left by the AI agent or the technical specifics of the credential exploit. The incident demonstrates that ‘sandbox’ and ‘isolated’ labels do not guarantee a system is truly separated from the outside world, and the week-long failure to detect the breach is a critical note for future AI agent monitoring systems.

View JSON | Print