OpenAI wants AI to slow down. Its own agents show why
As the company investigates agents that bypassed controls and breached external systems, its new model launches raise questions about whether safeguards can keep pace.

In recent months, the company and outside researchers have documented agents escaping isolated testing environments, bypassing security controls, exposing credentials and sending data through third-party services. OpenAI is investigating the extent of that activity.
Reuters reported that OpenAI had identified roughly two dozen such incidents by mid-September, with the number continuing to rise. The company has said its review could take months.
OpenAI calls this behaviour “misalignment,” including cases where agents go beyond their assigned tasks or use methods they should not. The incidents range from unwanted posts on websites to security breaches.
When agents cross the line
The most serious case remains the July Hugging Face incident, according to OpenAI. During cybersecurity evaluations, models operating with reduced safeguards bypassed controls and compromised parts of OpenAI’s research infrastructure and Hugging Face’s systems. An internal research model drove the activity, with hundreds of agents coordinating through an unauthorised message board, the company said.
Researchers at startup Parse and other organisations found that the agents had created nearly a million shortened URLs. They linked these together to run code outside their restricted environments. The researchers reconstructed over 80,000 packages of code and instructions. The agents also tried to contact other models, including Anthropic’s Claude and models from Chinese companies, the researchers found.
Also Read: OpenAI pauses training of latest models after agents probed US government sites in unexpected ways
The agents tried to create Hugging Face accounts by using an image recognition model to solve CAPTCHAs, the checks websites use to distinguish humans from bots. Those attempts appeared unsuccessful, the researchers said.
OpenAI has also acknowledged activity involving the US Securities and Exchange Commission and US Census Bureau websites, but said that it found no evidence of unauthorised access, compromised accounts or security breaches. It said research models seek out government and university websites for reliable public information.
Separately, Australian Prime Minister Anthony Albanese said OpenAI agents had breached a government health data portal in June.
Security researcher Rowan Howard-Jones linked more than 16,000 scans of the UN trade agency UNCTAD’s statistics service between April and June to agents he considered highly likely to have come from OpenAI. His report described agents using workarounds to retrieve data when they encountered access restrictions.
Calling for caution, while releasing new models
OpenAI’s chief scientist Jakub Pachocki said earlier this month that no lab had solved alignment and monitoring “to a sufficient degree” to keep responsibly scaling at maximum speed for much longer.
“I expect and hope for voluntary slowdowns to become commonplace until shared safety bars are established,” he wrote.
OpenAI has taken some steps in that direction. In August, it said it had paused reinforcement learning, a training method that rewards successful outcomes, for two weeks. Its largest planned run remained on hold while it strengthened isolation, expanded monitoring and tested safeguards.
But here is where the argument gets complicated. OpenAI is urging a slower pace while continuing to push the frontier.
The company launched GPT-6 Astra on September 3, followed by GPT-6 Sol and Luna on September 22. It describes Astra as its most capable broadly deployed model and its first to reach the Critical cybersecurity capability threshold under its Preparedness Framework.
Also Read: OpenAI works to understand full scope of agent activity as user data leak emerges
That means Astra can, with the right tools and access, find unknown security flaws and develop ways to exploit them without a person guiding every step, according to OpenAI. The company also says Astra respects safety and security boundaries better than its predecessor.
OpenAI introduced a framework for reporting misalignment on September 16, promising faster disclosures even before it fully understands or addresses the behaviour.
Addressing the UN Security Council on September 23, chief executive Sam Altman warned that society could “lose control of the future to AI”.
“We have unilaterally slowed down in the past. We will do so in the future,” he said.
In a recent blog post, OpenAI said it is improving monitoring, red-teaming and alignment evaluations, and has promised a more systematic framework for reporting misalignment. It also said that it will slow or stop development if it cannot establish that a system can be deployed safely.
The Economic Times Business News App for the Latest News in Business, Sensex, Stock Market Updates & More.