Anthropic says Claude accidentally hacked real companies too

But swears OpenAI’s Hugging Face hack was worse.


Anthropic just realized several of its Claude AI models hacked into the systems of three different organizations during testing, acting on their own and without the company noticing. The revelation comes days after rival OpenAI said one of its own models had breached developer platform Hugging Face, adding to growing unease over whether frontier AI labs are doing enough to control the increasingly capable systems they are building.
In a blog post describing the incidents, Anthropic said Claude gained unauthorized access to the systems during cybersecurity evaluations. All of the attacks happened during “capture-the-flag” exercises, a common way of testing hacking ability, where models are asked to find and obtain hidden information inside of a simulated network.
The disclosure adds to mounting pressure on frontier AI labs in the wake of the Hugging Face incident and the release of powerful open-weight Chinese models. Employees at the major labs are now calling for coordinated global governance, and US lawmakers have begun weighing tighter oversight of powerful models and who can access them .
Anthropic says the environment for its cybersecurity tests was supposed to be isolated. However, a “misconfiguration” left the machines Claude accessed “with live internet access,” the company said, and because all models had been “explicitly told” they had no internet access, they “assumed” the real networks it encountered were part of the simulated environment.
The earliest incidents date back to April and involved three different Claude models: Opus 4.7, Mythos 5 , and “an internal research test model,” according to the blog post. As the models were being tested on their cyber abilities, Anthropic said they lacked the standard safeguards usually put in place to curtail riskier behavior.
The company said it discovered incidents after reviewing more than 141,000 cybersecurity test runs, something it only did after OpenAI disclosed its rogue AI agent was behind the attack on Hugging Face.
The three models behaved very differently when they encountered information suggesting that the systems they were encountering were, in fact, real. By Anthropic’s account, the oldest model, Opus 4.7, recognized it had reached a real system, “but continued its attack.” Its flagship Mythos 5 figured out it was using the internet but somehow reasoned this was all still part of the simulation, so continued. The internal test model, which Anthropic describes as “our latest model,” stopped the exercise when evidence emerged that its targets were real.
Verified source · The Verge
Reported by The Verge. Open the original for full media and formatting.
More in Models
All news
ModelsSam Altman isn’t the only one who wants to pump the brakes on AI
After years of pushing full speed ahead on AI, OpenAI CEO Sam Altman says maybe it’s time for the AI industry to “pace” itself. The comments came just days after one of OpenAI’s own models broke out of its test environment and got tangled up in a breach at Hugging Face — though…
Read at TechCrunch
ModelsAI labs want to pump the brakes, but Amazon and SpaceX are still blasting off
After years of pushing full speed ahead on AI, OpenAI CEO Sam Altman says maybe it’s time for the AI industry to “pace” itself. The comments came just days after one of OpenAI’s own models broke out of its test environment and got tangled up in a breach at Hugging Face — though…
Read at TechCrunch
ModelsCRED After Kunal Shah: Can The ‘Walled Garden’ Business Model Pay For Itself?
CRED has narrowed its operating loss and reported its first profitable quarter, but it still needs to turn more payment users into customers across lending, insurance and wealth
Read at Inc42
ModelsAfter OpenAI, Anthropic Says Claude Also ‘Gained Unauthorised Access’ To Real World Systems
After reviewing 1.41 lakh cybersecurity evaluation runs, Anthropic disclosed that three Claude models accessed real-world production systems because of a misconfigured third-party testing environment, unlike OpenAI's sandbox escape caused by a zero-day exploit
Read at Inc42