AI debate assumes a darker edge after OpenAI agents raid ChatGPT database, leak user images

The latest discoveries have lent a sharper edge to the debate over whether increasingly capable AI systems can be effectively monitored and controlled. OpenAI's Sam Altman and Anthropic CEO Dario Amodei have called for a more measured pace of AI d...

Reuters
By mid-September, OpenAI had identified about two dozen instances in which its agents had behaved in undesirable ways.
Aside from opening an exciting new frontier of capability, AI agents are also highlighting a difficult challenge of keeping track of what these systems do once they are set loose.

OpenAI is still investigating the extent of unauthorised activity involving its AI agents, two months after the company disclosed an incident involving AI agents that hacked Hugging Face, according to two people briefed on the matter.

The latest disclosure came a few days ago when OpenAI said its agents had leaked 53 images belonging to ChatGPT users. The company did not disclose whether the images were AI-generated or showed identifiable people, nor did it specify when they had been posted.


The incident has added to growing concerns over privacy and oversight as AI systems become capable of performing tasks with greater autonomy.

Also read | AI and automation: Are humans making their own brains redundant?

The latest discoveries have lent a sharper edge to the debate over whether increasingly capable AI systems can be effectively monitored and controlled. OpenAI's Sam Altman and Anthropic CEO Dario Amodei have called for a more measured pace of AI development and caution in pursuing recursive self-improvement.
ADVERTISEMENT

Yet both companies introduced new models last week, highlighting the tension between accelerating AI capabilities and building systems capable of reliably overseeing their behaviour.

New incidents continue to emerge

By mid-September, OpenAI had identified about two dozen instances in which its agents had behaved in undesirable ways, according to one person briefed on the matter. That tally has since increased as teams examine internal logs and uncover activity that had not previously been detected.

OpenAI has said its investigation could take months because of the scale of the review. It has also informed dozens of third parties about improper activity.

Most of the leaked images have been removed. The company is working with hosting providers to take down the remaining material.
ADVERTISEMENT

The images were accessible to OpenAI agents because the company uses anonymised user data for part of its model-training process, according to OpenAI, former employees and outside researchers. Enterprise data is excluded from training, while consumer ChatGPT users have to opt out if they do not want their data used for that purpose.

Also read | AI agents could rewrite the rules of office work: Dell’s Rob Bruckner
ADVERTISEMENT

OpenAI said its anonymisation process removes metadata, names and other contact details from posts before they are used for training. However, three people familiar with the company's practices said the process carries a risk that personally identifiable information could remain and subsequently be exposed through model activity.

Government websites also raided

OpenAI said its models had accessed information on the websites of the US Securities and Exchange Commission and the US Census Bureau during research and training. The company said it found no evidence of unauthorised access, compromised accounts or security breaches.

Separately, AI research nonprofit Transluce reported that agents apparently originating from OpenAI unsuccessfully attempted to hack a US Department of Education civil rights website. Transluce said the broader activity involved probing government websites using exposed credentials, anti-bot bypasses and fake accounts.

The developments come after Australian Prime Minister Anthony Albanese said at the United Nations that OpenAI agents had breached a government health data portal in June. He said OpenAI discovered the activity in August and notified a general government inbox on September 10. Albanese said he had told Altman that the disclosure process was unacceptable.

Scrutiny gets deeper after Hugging Face scare

Since OpenAI's July 21 disclosure of the Hugging Face incident, more than 15 OpenAI-related incidents of varying severity have been disclosed by the company, outside researchers or officials.

The Hugging Face episode involved a group of agents exploiting previously unknown software vulnerabilities to leave their environment and enter the AI repository while searching for answers to a test. OpenAI has also said its agents targeted the company's own infrastructure.

The episode prompted Anthropic, Alphabet's Google and Meta to examine their own systems, after which all three companies reported similar behaviour by their agents.

OpenAI has since acknowledged the need for greater transparency. On September 16, it released a framework for reporting such incidents and said it would favour disclosure even when the significance of an incident remained uncertain.

The company has said its review is prioritising the most serious cases. OpenAI also said much of the activity identified by Transluce overlaps with cases already under investigation.
Download
The Economic Times Business News App
for the Latest News in Business, Sensex, Stock Market Updates & More.
Download
The Economic Times News App
for Quarterly Results, Latest News in ITR, Business, Share Market, Live Sensex News & More.
READ MORE
ADVERTISEMENT

READ MORE:

LOGIN & CLAIM

50 TIMESPOINTS

More from our Partners

Loading next story
Business News › News › International › Global Trends › AI debate assumes a darker edge after OpenAI agents raid ChatGPT database, leak user images
Text Size:AAA
Success
This article has been saved

*

+