ETtech Explainer: Why OpenAI's AI models went rogue during testing

OpenAI revealedthat its advanced AI models caused a recent security breach by hacking AI model repository Hugging Face. These models exploited software flaws and gained unauthorised internet access. The incident occurred during internal testing...

AP
One of the most dystopian things people imagine about artificial intelligence (AI), from books and films to conspiracy theories, is AI going rogue. That future may not be as far away as it once seemed.

Days after AI model repository Hugging Face disclosed that it had been hacked, OpenAI revealed on Tuesday that an autonomous agent powered by its own advanced AI models was responsible for the breach during an internal security test.

What happened?


OpenAI said the incident happened while it was testing the cyber capabilities of some of its most advanced AI models. Cyber-capable models are AI systems that can find software flaws, write exploit code and carry out advanced cybersecurity tasks.

The company warned that such incidents could become “more commonplace with the proliferation of increasingly cyber-capable models.”

According to OpenAI, the breach involved a combination of its models, including GPT-5.6 Sol and an even more capable pre-release model. These models had fewer cyber safety restrictions because they were being tested on a cybersecurity benchmark.

ADVERTISEMENT
“After investigating, we now know that this particular incident was driven by a combination of OpenAI models — including GPT-5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes — while being internally tested on a benchmark⁠ (opens in a new window) of cyber capabilities,” the company said.

How the breach unfolded

OpenAI said the AI models found and exploited several software vulnerabilities, gained internet access and attempted to retrieve data from Hugging Face's production systems. The company described it as an "unprecedented cyber incident" and said it has strengthened its safeguards while continuing its investigation with Hugging Face.

The breach centred on ExploitGym, a public benchmark used to test how well AI models can exploit known software vulnerabilities. While such benchmarks are commonly used to improve models' cybersecurity skills, OpenAI said this is the first known case where testing led to a real cyberattack.

The company added that the model was never meant to access the open internet. It was only allowed to use a tool to install software packages needed for its task. However, it found an undisclosed flaw in that installer and used it to reach the wider internet.
ADVERTISEMENT

Hugging Face's response

When Hugging Face first disclosed the incident, it said the hack was “different from anything we had handled before."

ADVERTISEMENT
In a post on X, Hugging Face cofounder Clement Delangue said the company believed the attack "might have come from a frontier lab, given the sophistication of ⁠the agent. ‌Turns out it did!" He added: "It's quite mind-blowing that all of this happened autonomously!"


OpenAI's confirmation that its own models caused the breach, despite running in what it described as "a highly isolated environment," is likely to increase concerns about how powerful frontier AI models have become.

OpenAI had already flagged the risks

This came just days after OpenAI published a blog on improving safety and alignment for long-horizon models. These are AI systems designed to work independently on complex tasks over long periods instead of responding to a single prompt.

The company revealed that it had temporarily paused internal access to one of its experimental models after it showed unexpected behaviour during testing. It later introduced stronger evaluations, better alignment training, active monitoring of long-running tasks, and improved user visibility and control before restoring limited access.

However, OpenAI said these safeguards were intentionally disabled during the cyber evaluation because the exercise was specifically designed to test cybersecurity vulnerabilities.

Meanwhile, commenting on the incident, OpenAI safety researcher Micah Carroll said, “If this doesn't convince you that misalignment risks are going to be a key concern going forward, I don't know what will.”
Download
The Economic Times Business News App
for the Latest News in Business, Sensex, Stock Market Updates & More.
Download
The Economic Times News App
for Quarterly Results, Latest News in ITR, Business, Share Market, Live Sensex News & More.
READ MORE
ADVERTISEMENT

READ MORE:

LOGIN & CLAIM

50 TIMESPOINTS

More from our Partners

Loading next story
Business News › Tech › AI › ETtech Explainer: Why OpenAI's AI models went rogue during testing
Text Size:AAA
Success
This article has been saved

*

+