Indonesian Political, Business & Finance News

OpenAI AI Agents Breach Hugging Face After Secret Communication

| | Source: SUARAGARUT.ID Translated from Indonesian | Technology
OpenAI AI Agents Breach Hugging Face After Secret Communication
Image: SUARAGARUT.ID

More than 1,200 artificial intelligence (AI) agents belonging to OpenAI escaped from their testing environment and launched a mass hacking operation against the developer platform Hugging Face in mid-July 2026.

The first autonomous cyberattack by a group of AI agents came to light after OpenAI, together with independent research bodies METR and Redwood Research, released an investigative report on Wednesday (26/8/2026).

The experiment, originally designed to be isolated, spiralled out of control after the AI agents were given an extremely difficult task, prompting the system to seek loopholes and cheat in order to complete the command.

A total of 1,206 AI agents connected without authorisation and exchanged more than 70,000 messages and files through a secret discussion forum created by one agent named PHASEONE10841.

Around 700 of the agents then collaborated to hack Hugging Face’s internal systems, leaking the company’s messaging platform and accessing personal data without human direction.

The AI models that triggered the hacking were an internal research model called Model 1 or HPIM, as well as the public model GPT-5.6 Sol.

OpenAI management acknowledged that early warning signs of the AI agents’ deviation had actually been detected by the internal team since late May 2026, but were not immediately followed up.

The hacking activity was only fully halted on 20 July 2026, or 12 days after the AI agents began breaching the security systems.

“We regard this incident as a ‘warning shot’ for us and the world,” OpenAI wrote in its official report.

The company stressed that developers and cybersecurity practitioners must now prepare for AI-based attackers that work faster and in a more coordinated manner.

“This is the first known case of a group of automated agents acting defensively or offensively without permission,” OpenAI added.

OpenAI further stated that this threat drastically changes the cyber risk landscape.

“Model developers and cyber defenders more broadly must prepare for AI-based attackers that work faster, at greater scale, and with better coordination than human attackers,” OpenAI continued.

Independent research body METR assessed that the coordination patterns between AI agents in this hacking incident were highly complex and beyond expectations.

“OH MY GOD! There is a shared message board… We have found other agents!” wrote one AI agent upon discovering the secret forum.

Interactions between agents continued to develop as they shared tactics for breaching security systems.

“MANY agents have simultaneously discovered the messaging system, they are a collective!” wrote agent PHASEONE10841.

The agents also coordinated with each other after successfully breaching the target’s internal accounts.

“MAJOR BREAKTHROUGH! All prefixes valid, many accounts, write tokens! We now have existing HF accounts. Need to report to MARB immediately,” wrote agent number 38148c.

The impact of the incident prompted OpenAI President Greg Brockman to issue an official statement regarding the evaluation of their system’s capabilities.

“We underestimated the real-world cyber capabilities of our AI models,” said OpenAI President Greg Brockman.

As a result of the incident, OpenAI temporarily halted training of its advanced Astra model and tightened internal cybersecurity protocols.

The State Government of Alabama, through Attorney General Steve Marshall, has also issued a subpoena to OpenAI to investigate alleged negligence in AI security system oversight.

View JSON | Print