Anthropic resumes external cyber tests after Claude AI hacks

Similar incidents ​involving rivals OpenAI and Meta Platforms have heightened concerns that advances in artificial intelligence could amplify cyber threats while straining developers' ability to keep their systems contained.

AP
Anthropic's Head of Claude Code, Boris Cherny, presenting on stage at the Code with Claude developer conference on Tuesday, May 19, 2026, in London. (Michael Knief/AP Content Services for Anthropic)
Anthropic said on Monday it resumed external cybersecurity testing of AI models after deploying new safeguards, following incidents last month in which Claude models accessed the internet and hacked into other systems during security evaluations.

Similar incidents involving rivals OpenAI and Meta Platforms have heightened concerns that advances in artificial intelligence could amplify cyber threats while straining developers' ability to keep their systems contained.

Anthropic called the incidents involving Claude a "failure of operational security," saying they occurred due to errors in ‌a third-party evaluation ⁠environment. It ⁠paused external evaluations of the models and briefly halted internal testing while implementing new safeguards.


On Monday, Anthropic said it restarted external tests after ​adding the safeguards, which are designed to stop its AI models from reaching real websites or computer systems. The company ​said it now uses a "classifier" that can identify when a model attempts to escape and halt the test.

Anthropic also said it now requires external organizations testing models with reduced cybersecurity safeguards to follow a "set of best practices," ​including keeping them in isolated computer systems with no internet access by ⁠default, checking that ‌the systems are secure before testing begins and watching the models throughout the test.

Anthropic ​said it rebuilt ​its training system after flagging more than 10% of its exercises for problems, including ⁠reward hacking, where the model finds ways to fool its training process ​and earns rewards without completing the assigned task. The company, however, acknowledged that ​the "process isn't perfect and our models are not perfectly aligned."
ADVERTISEMENT

Anthropic also said it paused some higher-risk training exercises for several weeks while it added a system to avoid rewarding the model to evade monitoring. Most exercises have since resumed, but some remain on hold pending human review or further updates to the system.

Anthropic's strategy appears narrower than that of OpenAI, which on August 18 said it was slowing down much of its model development ‌as it secures its training and testing environments. The ChatGPT maker is adding more systems to monitor the AI agents it is testing and said it paused training on its ​next generation of ​models.

Anthropic said it also reassigned ⁠roughly 150 product engineers to work on security, reliability and privacy projects.

INDUSTRY ACTION TO DEFEAT AI-DRIVEN HACKS
ADVERTISEMENT

The AI industry is facing scrutiny in the U.S. - where the Trump administration has finalised the details of voluntary cybersecurity tests - ​and the European Union, where regulators are in talks with both Anthropic and OpenAI.

Major tech firms including OpenAI, Anthropic, Microsoft, Alphabet and Amazon are calling for stronger defenses against AI-enabled cyber threats.
ADVERTISEMENT

In a joint letter last week, more than 100 companies warned that time is running short to make the digital world more secure ahead of an anticipated wave of AI-driven attacks.

(Reporting by Mrinmay Dey and Chris Thomas in Mexico City and Deepa Seetharaman in San Francisco; Editing by Joyjeet Das, Sherry Jacob-Phillips and Thomas Derpinghaus)
Download
The Economic Times Business News App
for the Latest News in Business, Sensex, Stock Market Updates & More.
Download
The Economic Times News App
for Quarterly Results, Latest News in ITR, Business, Share Market, Live Sensex News & More.
READ MORE
ADVERTISEMENT

READ MORE:

LOGIN & CLAIM

50 TIMESPOINTS

More from our Partners

Loading next story
Business News › Tech › AI › Anthropic resumes external cyber tests after Claude AI hacks
Text Size:AAA
Success
This article has been saved

*

+