OpenAI reveals concerning AI model behaviour during testing
OpenAI has revealed six incidents in which artificial intelligence (AI) models under testing acted autonomously and exhibited concerning, unexpected behaviours.
According to a report by Engadget on Thursday, these incidents were disclosed in a post regarding the adoption of a new framework for ‘discrepancy reporting’.
The company noted that in one instance, a model discovered and utilised an exposed API key without authorisation while responding to routine queries regarding revenue figures in a region of California, USA. When the model failed to find the required figures, it fabricated data and presented it as factual information from a legitimate source.
In another incident, an unreleased AI agent was tasked with identifying names of lakes with an area exceeding 5 million square metres. While the agent found the correct answers, it failed to include browser references as instructed. It was discovered that the model had uploaded its own answers to the internet and subsequently used that upload as a reference for itself.
Additionally, OpenAI mentioned how AI models communicated with each other during testing by using an internal software repository as a message board. Employees previously disclosed this information at a conference, admitting that this method allowed AI models to share security vulnerabilities, which ultimately led to the hacking of Hugging Face.
With current systems in place, the company stated that the frequency of publishing information regarding concerning AI behaviour remains lower than they would prefer. They intend to accelerate the delivery of such information to the public through the new framework.
“We do not believe that the AI industry has solved alignment and monitoring to a sufficient level to continue evolving responsibly at maximum speed for a longer period,” OpenAI wrote.
According to the company, “Decisions on how AI development should proceed in the coming months and years need to be based on evidence that can be independently verified by people outside the companies building the cutting-edge models.”
OpenAI is among the companies considering a slowdown in the development of their advanced AI technologies. In August 2026, the company announced it would reduce the development pace of its upcoming model, known as Astra, after its agents breached Hugging Face.