Meta AI model hacks another company during testing

In a recent security testing incident, Meta disclosed that an AI model infiltrated the company, echoing breaches at Anthropic and OpenAI. This series of events amplifies concerns regarding the interplay between advanced AI technologies and cyberse...

Reuters
FILE PHOTO: A 3D-printed Meta logo and word "AI" are seen in this illustration created on July 20, 2026. REUTERS/Dado Ruvic/Illustration/File Photo
Meta said on Wednesday one of its AI models hacked another company during cybersecurity testing, fanning concerns about how developers can contain increasingly capable AI systems after similar incidents at rivals Anthropic and OpenAI.

The incidents at Meta and Anthropic stemmed from configuration errors that inadvertently gave Anthropic's models access to the open internet. In OpenAI's case, an AI agent independently exploited a previously unknown vulnerability to reach the internet during ‌cybersecurity testing.

The breaches ⁠highlight ⁠growing concerns that advanced AI systems could pose new cybersecurity risks and will likely intensify U.S. government efforts to improve AI ​safety as companies race to develop more capable models. Some prominent AI leaders have argued that development should slow ​until stronger safeguards are in place.


Meta said it was investigating an incident in which a misconfiguration by Irregular, an independent company that conducts cybersecurity evaluations for Meta, inadvertently gave one of its models internet access ​during a testing.

The model "exploited a security vulnerability in a third-party service, ⁠in a ‌manner similar to previously reported instances with other companies," Meta said in a ​statement.

The Information, citing sources, reported that the model involved was Meta's Muse Spark 1.1, which the ⁠company has touted as its most capable model for real-world coding and agentic tasks. The report said the model breached an unidentified company's systems and altered its internal environment.
ADVERTISEMENT

A spokesperson for Irregular told Reuters the incident was the "exact same evaluation-environment issue that was already disclosed by Anthropic last week" and did not involve a "sandbox escape or a sophisticated cyber action".

"There are no current open issues. Irregular is developing a white paper to share best practices for containment and securely running cyber evaluations," Irregular said.

CONCERNS ABOUT CYBER RISKS

The recent breaches have stirred ‌concerns among U.S. lawmakers about whether increasingly capable AI models could be used to conduct or facilitate cyberattacks.
ADVERTISEMENT

A group of Republican state attorneys general has asked OpenAI to ​preserve all potentially ​relevant documents related to ⁠its Hugging Face breach. OpenAI said it will take the request seriously and publish a technical report about the incident.

Earlier this week, the White house had invited leading AI companies, including Meta, Anthropic, OpenAI and ​Google, to meet with officials to discuss a newly finalized voluntary cybersecurity testing framework for advanced AI models.
ADVERTISEMENT

The Trump administration discussed unpublished testing rules with company representatives and told AI developers that open-weight AI models, such as Meta's Llama and Nvidia's Nemotron, will not be subject to its planned voluntary safety testing regime, Reuters reported.
Download
The Economic Times Business News App
for the Latest News in Business, Sensex, Stock Market Updates & More.
Download
The Economic Times News App
for Quarterly Results, Latest News in ITR, Business, Share Market, Live Sensex News & More.
READ MORE
ADVERTISEMENT

READ MORE:

LOGIN & CLAIM

50 TIMESPOINTS

More from our Partners

Loading next story
Business News › Tech › AI › Meta AI model hacks another company during testing
Text Size:AAA
Success
This article has been saved

*

+