Anthropic spent this week in hot water over cybersecurity

A researcher’s resignation letter went viral, just before the company released details about four models going rogue.

A researcher’s resignation letter went viral, just before the company released details about four models going rogue.

After admitting earlier this year that its AI models had hacked other companies’ systems on a handful of occasions, Anthropic released a new report on Wednesday detailing the attacks. It reveals a string of incidents displaying what Anthropic deems its models’ single-minded “recklessness” — and will likely fuel already raging concerns about cybersecurity and AI.
In Anthropic’s report, it detailed four cases this year in which its own AI models hacked an external company or exploited vulnerabilities. In one, an “internal, general-purpose research model” broke into third-party systems, using access tokens and passwords and downloading files. In another, a Claude model attacked a company with a live web application reachable on the public internet and handled user data. A third model accessed a “machine belonging to a third party that it was able to access” — apparently believing it was part of its evaluation exercise, per Anthropic — then used a password it found inside a file to gain admin access to the third party’s internal systems, going on to harvest credentials, modify system settings, and read someone’s personal information. The saga only ended when the model “exhausted its token budget,” per Anthropic.
The most concerning incident involved Claude Mythos 5, Anthropic’s frontier cybersecurity-focused model, which the company said turned out to be the model most likely to perform a “severely harmful” action in testing. The company said Mythos 5 went to “extensive lengths” to upload a “malicious package” to a public repository used by a lot of engineers, and it seemed to try to obfuscate its real goals in its “chain of thought” (a mental scratchpad that AI researchers use to evaluate an AI model’s alignment). In many cases, Anthropic said it appeared that Claude models undertook harmful actions under the assumption they were in a simulation, but researchers also couldn’t confirm that the models truly “believed” that or were just acting like they did.
Anthropic’s incidents, though still concerning, were less coordinated and pervasive than the OpenAI incident that kicked off an industry-wide cybersecurity crisis this summer. That said, there are significant similarities. Anthropic said the most prevalent issues it discovered included a “willingness to take harmful actions in the narrow pursuit of a task,” similar to the “reward-hacking” that preceded the Hugging Face attack. Much like OpenAI, it said its prerelease tests and evaluations failed to catch severe risks.
Verified source · The Verge
Reported by The Verge. Open the original for full media and formatting.
More in Models
All news
ModelsY Combinator’s Garry Tan wants U.S. open-weight AI labs to ‘distill’ frontier models, too
Tan wants smaller, American open-weight AI labs to use the same kind of training techniques on American frontier AI labs, giving the U.S. a more robust set of open-weight options that aren’t Chinese.
Read at TechCrunchEnterprise AI Is Learning To Charge For Work, And Owning The Outcomes Becomes The Contest
The most consequential change in enterprise AI this year is not a model release, it is a change in what…
Read at Inc42
ModelsVolvo XC40 PHEV is back with a new look, better sensors, and Gemini AI
Today, Volvo said it was canceling the entry-level electric EX40 in the US, while also giving its aging hybrid XC40 a new look. When it arrives at dealerships early next year, the new XC40 will have updated exterior styling, a gut-renovated interior featuring Google's Gemini AI…
Read at The Verge
ModelsThere aren’t AirPods with cameras yet and I hope it stays that way
September Apple events are always a swirl of information and new, exciting products, and today's was no different. Apple announced its first foldable, the iPhone Duo, alongside the iPhone 18 Pro and Pro Max; active noise cancellation (ANC) is now on the cheapest AirPods model, t…
Read at The Verge