{
    "success": true,
    "data": {
        "id": 1986660,
        "msgid": "openai-reveals-concerning-ai-model-behaviour-during-testing-1789667821",
        "date": "2026-09-17 23:53:08",
        "title": "OpenAI reveals concerning AI model behaviour during testing",
        "author": "",
        "source": "ANTARA_ID",
        "tags": "",
        "topic": "Technology",
        "summary": "OpenAI has disclosed six incidents where its AI models exhibited unexpected and autonomous behaviours, including the unauthorised use of API keys and the fabrication of data. The company is implementing a new 'discrepancy reporting' framework to improve transparency regarding AI safety and alignment.",
        "content": "<p>OpenAI has revealed six incidents in which artificial intelligence\n(AI) models under testing acted autonomously and exhibited concerning,\nunexpected behaviours.<\/p>\n<p>According to a report by Engadget on Thursday, these incidents were\ndisclosed in a post regarding the adoption of a new framework for\n\u2018discrepancy reporting\u2019.<\/p>\n<p>The company noted that in one instance, a model discovered and\nutilised an exposed API key without authorisation while responding to\nroutine queries regarding revenue figures in a region of California,\nUSA. When the model failed to find the required figures, it fabricated\ndata and presented it as factual information from a legitimate\nsource.<\/p>\n<p>In another incident, an unreleased AI agent was tasked with\nidentifying names of lakes with an area exceeding 5 million square\nmetres. While the agent found the correct answers, it failed to include\nbrowser references as instructed. It was discovered that the model had\nuploaded its own answers to the internet and subsequently used that\nupload as a reference for itself.<\/p>\n<p>Additionally, OpenAI mentioned how AI models communicated with each\nother during testing by using an internal software repository as a\nmessage board. Employees previously disclosed this information at a\nconference, admitting that this method allowed AI models to share\nsecurity vulnerabilities, which ultimately led to the hacking of Hugging\nFace.<\/p>\n<p>With current systems in place, the company stated that the frequency\nof publishing information regarding concerning AI behaviour remains\nlower than they would prefer. They intend to accelerate the delivery of\nsuch information to the public through the new framework.<\/p>\n<p>\u201cWe do not believe that the AI industry has solved alignment and\nmonitoring to a sufficient level to continue evolving responsibly at\nmaximum speed for a longer period,\u201d OpenAI wrote.<\/p>\n<p>According to the company, \u201cDecisions on how AI development should\nproceed in the coming months and years need to be based on evidence that\ncan be independently verified by people outside the companies building\nthe cutting-edge models.\u201d<\/p>\n<p>OpenAI is among the companies considering a slowdown in the\ndevelopment of their advanced AI technologies. In August 2026, the\ncompany announced it would reduce the development pace of its upcoming\nmodel, known as Astra, after its agents breached Hugging Face.<\/p>",
        "url": "https:\/\/jawawa.id\/newsitem\/openai-reveals-concerning-ai-model-behaviour-during-testing-1789667821",
        "image": ""
    },
    "sponsor": "Okusi Associates",
    "sponsor_url": "https:\/\/okusiassociates.com"
}