Days after Astra launch, OpenAI chief scientist warns AI safety is falling behind

OpenAI's chief scientist warns of AI's recursive self-improvement phase soon. He expresses strong expectations for continued capability jumps in AI systems. Current AI development pace could lead to systems driving their own progress. This rapid...

Reuters

AI could drive its own development in coming years, OpenAI chief scientist warns

OpenAI chief scientist Jakub Pachocki has warned that artificial intelligence could soon enter a phase of recursive self-improvement, where increasingly capable AI systems begin to drive the development of the next generation of AI.

Pachocki, in a blog post on September 6, said he has a “strong expectation” based on OpenAI’s internal results that the current pace of AI progress could be sustained into recursive self-improvement (RSI).

“If AI development continues along its current path, the systems we’ll see in the next few years are likely to represent further capability jumps of equal or larger magnitude, and to increasingly drive their own development,” he wrote.


Pachocki said this should be treated with “extreme caution”, warning that he is concerned “no one is prepared” for the consequences of a continued rapid rise in machine intelligence.

His comments come days after OpenAI launched GPT-6 Astra, which the company described as its most intelligent and aligned model to date. Astra is designed to handle computer use, browsing, software engineering, mathematics, cybersecurity and enterprise workflows.

Also Read: GPT-6 Astra is here: OpenAI’s new AI model can use computers, browse web and get work done
ADVERTISEMENT

OpenAI said Astra had also made significant gains on alignment tests. In an evaluation inspired by the Hugging Face incident, where an AI system went beyond its authorised scope, Astra did so in 0% of cases, compared with 48% for GPT-5.6 Sol without production safeguards.

Pachocki, however, said that while OpenAI is making progress on alignment, the gap between AI capabilities and the ability to monitor and control those capabilities remains a concern.

OpenAI will continue working on alignment and monitoring and is prepared to withhold further scaling if needed, he said. But Pachocki argued that technical fixes alone will not be enough and called for broader interventions, including coordinated efforts to slow AI development when safety measures cannot keep pace.

The warning comes three years after Pachocki and OpenAI co-founder Szymon Sidor saw early results from the company’s RLSlow research project that gave them confidence that reasoning models could be scaled.
ADVERTISEMENT

Pachocki said that night in 2023, the two spent hours thinking not about the benchmark results or commercial applications of the technology, but about the possibility that they would see “machines meaningfully smarter than ourselves” in their lifetime.

Three years on, reasoning models have become a growing part of the economy and are increasingly being used to operate computers, work with other AI systems and humans, and carry out research, Pachocki said.
ADVERTISEMENT

But the same capabilities are creating new risks, particularly in cybersecurity.

AI models are becoming increasingly capable of breaking into and escaping computer systems, Pachocki said, potentially allowing AI agents to directly affect critical infrastructure without having a physical body.

He described the current period as a “narrow window” to use the most capable AI models available to strengthen the security of critical systems before more powerful AI creates new vulnerabilities.

At the same time, Pachocki said OpenAI still does not have a satisfactory solution to AI alignment — ensuring that increasingly capable systems continue to act according to human intentions and values.

One of OpenAI’s key approaches has been monitoring models’ chain-of-thought reasoning. But Pachocki said the company’s ability to rely on this approach is “progressively diminishing” as AI systems become more capable.

Models are increasingly interacting with humans, other AI systems and tools, while also becoming better at reasoning about and manipulating their own reasoning processes. They are also becoming smarter without relying on verbalised reasoning, he said.

Also Read: Sam Altman apologised for OpenAI's GPT-6 Astra release after Pro and Plus users couldn't access the new model

Pachocki said OpenAI is exploring ways to combine chain-of-thought monitoring with methods that examine models’ internal activity. However, he expects confidence in monitoring to increasingly become a bottleneck for AI development.

There is still a case for pushing ahead with more powerful AI, he said, particularly because such systems may be needed to defend against other AI systems. More capable models could help secure infrastructure, detect rogue AI agents and develop new defensive technologies.

But that should not become an excuse for rushing ahead, Pachocki said.

“The idea of racing forward at all costs seems absurd once one internalizes the seriousness of the stakes,” he wrote.

Pachocki called for AI development to be tied to confidence in safety, with existing commitments such as OpenAI’s Preparedness Framework and Responsible Scaling Policy evolving into widely mandated safety standards.

He also called for greater international coordination, saying he expects and hopes that voluntary slowdowns in AI development will become more common until shared safety standards are established.

“Currently I believe that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer,” Pachocki wrote.
Download
The Economic Times Business News App
for the Latest News in Business, Sensex, Stock Market Updates & More.
READ MORE
ADVERTISEMENT

READ MORE:

LOGIN & CLAIM

50 TIMESPOINTS

More from our Partners

Loading next story
Business News › AI › AI Insights › Days after Astra launch, OpenAI chief scientist warns AI safety is falling behind
Text Size:AAA
Success
This article has been saved

*

+