OpenAI finds evidence other AI agents escaped containment as it widens hacking probe
The discovery of additional rogue behaviour at OpenAI, even if limited in nature, could feed growing appetite for regulation coming out of the White House and elsewhere. The expanded investigation by OpenAI was launched shortly before its primary...

The new breakouts were uncovered during the company's publicly announced investigation into how one of its agents escaped what was meant to be a contained testing environment this month, the two people said, and OpenAI is now looking into those instances as well. One of the sources said that the escapes were limited in nature and that none of the agents were thought to have left OpenAI's network.
An OpenAI spokesperson referred to a statement issued by the company on Tuesday that said it was reviewing "broader activity from our models" in addition to the Hugging Face intrusion.
The discovery of additional rogue behaviour at OpenAI, even if limited in nature, could feed growing appetite for regulation coming out of the White House and elsewhere. The expanded investigation by OpenAI was launched shortly before its primary rival, Anthropic, disclosed that its models were also responsible for a series of break-ins that led to breaches at three other companies dating back to April, according to the two sources and a third source familiar with the matter. The recent discovery of other past breakouts at OpenAI has not previously been reported.
Growing concerns
AI safety experts said the new disclosures paint a portrait of a group of cutting-edge labs whose ability to develop dangerous autonomous hacking agents outstrips their ability to keep them under control.
Reuters could not establish exactly how many incidents OpenAI investigators found or the timings or circumstances under which they occurred. The three sources said OpenAI and outside experts were examining log data from earlier in the year in a bid to understand what took place.
OpenAI first launched the investigation following the early July intrusion at Hugging Face, where an OpenAI agent went haywire for days inside another company's network in a botched effort to cheat on an internal test.
As part of that hacking spree, OpenAI said that four accounts at four other companies were also compromised. One of those companies was New York-based Modal, corporate officials there said. Chiodo said his concerns were heightened by indications that neither OpenAI nor Anthropic were watching the agents as they went rogue.
Chiodo said that pointed to a lack of proper scrutiny.
Anthropic said that while it did have real-time monitoring in place, that monitoring had not been used "for this threat surface" due to a misunderstanding between the AI company and a partner.
Also read: AI on the loose: Why ChatGPT, Claude models went rogue and what happens next
New government oversight
The rapidly widening scope of the runaway AI agents story already has heightened pressure from lawmakers and officials across the United States and Europe to push for new government oversight of the labs whose models power them. "We're looking at controls," U.S. President Donald Trump told reporters on Thursday. On Friday, the European Commission said it held talks with OpenAI and Anthropic over the hacking incidents.
Mark Warner, the top Democrat on the U.S. Senate Intelligence Committee, said on Friday that the Anthropic incident "tells me that legislatively we're correct to require mandatory capabilities testing of these advanced models."
The Economic Times Business News App for the Latest News in Business, Sensex, Stock Market Updates & More.
The Economic Times News App for Quarterly Results, Latest News in ITR, Business, Share Market, Live Sensex News & More.