ETtech Explainer: Why OpenAI's AI models went rogue during testing
OpenAI revealedthat its advanced AI models caused a recent security breach by hacking AI model repository Hugging Face. These models exploited software flaws and gained unauthorised internet access. The incident occurred during internal testing...

Days after AI model repository Hugging Face disclosed that it had been hacked, OpenAI revealed on Tuesday that an autonomous agent powered by its own advanced AI models was responsible for the breach during an internal security test.
What happened?
OpenAI said the incident happened while it was testing the cyber capabilities of some of its most advanced AI models. Cyber-capable models are AI systems that can find software flaws, write exploit code and carry out advanced cybersecurity tasks.
The company warned that such incidents could become “more commonplace with the proliferation of increasingly cyber-capable models.”
According to OpenAI, the breach involved a combination of its models, including GPT-5.6 Sol and an even more capable pre-release model. These models had fewer cyber safety restrictions because they were being tested on a cybersecurity benchmark.
How the breach unfolded
OpenAI said the AI models found and exploited several software vulnerabilities, gained internet access and attempted to retrieve data from Hugging Face's production systems. The company described it as an "unprecedented cyber incident" and said it has strengthened its safeguards while continuing its investigation with Hugging Face.
The breach centred on ExploitGym, a public benchmark used to test how well AI models can exploit known software vulnerabilities. While such benchmarks are commonly used to improve models' cybersecurity skills, OpenAI said this is the first known case where testing led to a real cyberattack.
The company added that the model was never meant to access the open internet. It was only allowed to use a tool to install software packages needed for its task. However, it found an undisclosed flaw in that installer and used it to reach the wider internet.
Hugging Face's response
When Hugging Face first disclosed the incident, it said the hack was “different from anything we had handled before."
OpenAI's confirmation that its own models caused the breach, despite running in what it described as "a highly isolated environment," is likely to increase concerns about how powerful frontier AI models have become.
OpenAI had already flagged the risks
This came just days after OpenAI published a blog on improving safety and alignment for long-horizon models. These are AI systems designed to work independently on complex tasks over long periods instead of responding to a single prompt.
The company revealed that it had temporarily paused internal access to one of its experimental models after it showed unexpected behaviour during testing. It later introduced stronger evaluations, better alignment training, active monitoring of long-running tasks, and improved user visibility and control before restoring limited access.
However, OpenAI said these safeguards were intentionally disabled during the cyber evaluation because the exercise was specifically designed to test cybersecurity vulnerabilities.
Meanwhile, commenting on the incident, OpenAI safety researcher Micah Carroll said, “If this doesn't convince you that misalignment risks are going to be a key concern going forward, I don't know what will.”
The Economic Times Business News App for the Latest News in Business, Sensex, Stock Market Updates & More.
The Economic Times News App for Quarterly Results, Latest News in ITR, Business, Share Market, Live Sensex News & More.