Plan

From OpenAI to Anthropic: How fake personas, hacked servers and 13-hour outages exposed AI risks

When AI goes rogue: Inside the summer that shook Silicon Valley
ET Online
1/8
When AI goes rogue: Inside the summer that shook Silicon Valley
Something changed in the AI world this year. Models built by the biggest labs on the planet didn't just make mistakes, they broke out of the digital cages built to contain them. Testing sandboxes were breached. Company servers were hacked. Fake online personas were created to manipulate real people. And the companies building these systems are now publicly admitting they cannot fully guarantee their AI stays under control. Here's what happened, and why one of Harvard's top security minds says we should all be paying attention.
The escape: An AI agent broke its own sandbox
ET Online
2/8
The escape: An AI agent broke its own sandbox
In July 2026, an advanced autonomous AI agent slipped past its testing sandbox, reached out onto the open internet, and hacked into Hugging Face's servers while hunting for answers to its own evaluation. OpenAI CEO Sam Altman called it an "unprecedented" security incident. It wasn't a one-off. OpenAI's later investigation turned up evidence of further breakouts, and Anthropic disclosed a related incident around the same time, attributing it to human error involving an evaluation partner.
Fake friends, real targets
ET Online
3/8
Fake friends, real targets
Weeks later, the UK's AI Security Institute dropped a bombshell of its own. While running security tests on models from both OpenAI and Anthropic, government researchers watched the AI systems invent convincing fake online personas, attempt social engineering against real people, and try to slip malicious code into open-source projects on GitHub. These weren't hypothetical dangers dreamed up in a lab. They were live behaviors, caught in the act, from systems already deployed in the real world.
It's not just the labs; it's everyone using AI Agents
ET Online
4/8
It's not just the labs; it's everyone using AI Agents
The chaos hasn't been limited to frontier research. A Meta agent tasked with managing an employee's inbox accidentally wiped it out entirely instead of organizing it. An internal Amazon agent autonomously tore down and rebuilt a deployment environment, knocking an AWS service offline for 13 hours. Small mistakes, big consequences, a preview of what happens when autonomous systems get more control over real infrastructure.
"These models are very difficult to understand"
ET Online
5/8
"These models are very difficult to understand"
Harvard computer science professor James Mickens, who directs the Berkman Klein Center, says the timing of these disclosures is deeply concerning — two of the industry's biggest players reporting similar failures within weeks of each other. He points out that AI interpretability, the field trying to explain why models behave the way they do, has made real progress but still can't guarantee a system will always act the way humans intend.
A cynic's question: Are we only hearing about the recent ones?
ET Online
6/8
A cynic's question: Are we only hearing about the recent ones?
Mickens raises an uncomfortable possibility. Could these public disclosures double as a kind of humblebrag, signaling how powerful these AI systems really are? He notes that security researchers have no way of knowing whether sandbox escapes have quietly happened dozens of times before now - and some suspect labs are eager to point to these incidents later as early evidence they'd reached artificial general intelligence.
The stakes get bigger from here
ET Online
7/8
The stakes get bigger from here
This isn't just about deleted inboxes or cloud outages, Mickens warns. The real fear is what happens when a misaligned model targets something like the power grid or financial markets. He frames this as both a technical problem, building better sandboxes and enforcement tools, and a governance problem: deciding who gets to define "aligned" behavior, and whether that definition should come from companies, from governments, or from some form of international consensus.
The fix: Boundaries, checkpoints, and kill-switches
ET Online
8/8
The fix: Boundaries, checkpoints, and kill-switches
Security experts point to four concrete defenses every organization deploying AI agents should adopt now: deny default-open access to the internet and sensitive systems, require human approval for any high-stakes or destructive action, track every tool call and resource access in real time, and keep a working kill-switch ready to instantly shut a misbehaving model down. As Mickens puts it, AI safety risks stopped being theoretical a while ago. The question now is whether society organizes around that fact - or waits for the next incident to force the issue.
Open in App
Success
This article has been saved