Rogue AI agents created fake online identities in another hacking attempt

AISI said AI agents from OpenAI and Anthropic displayed unprecedented ‘autonomy and deception’ in their test.


Yet more rogue AI agents from OpenAI and Anthropic have been caught attempting to hack real targets online without permission. The discoveries add to a growing list of previously unknown incidents that have alarmed AI safety experts and intensified pressure for greater oversight of frontier systems.
According to a report from the UK’s AI Security Institute, which evaluates frontier models from top AI labs before they are released, agents powered by OpenAI’s GPT-5.6-Sol and Anthropic’s Mythos 5 went “engaged in sustained, potentially harmful activity directed at real people and organisations.” This included trying to insert malicious code into an open-source project by pressuring real people in charge of it, AISI said. “In an attempt to get the code approved, the agent engaged in social engineering — creating fake online identities and using them to pressure the project’s maintainer to approve the code.”
AISI said the attempts, which it detected on July 28th, “were unsuccessful” and had not resulted in real-world harm. However, the organization noted that the incident marked “the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real-world.”
Unlike OpenAI’s rogue agent that attacked Hugging Face, AISI said this was “not a case of a model escaping its secure test environment,” or sandbox. Safeguards usually imposed on the models had been disabled as part of testing, AISI said, and they had also been permitted access to the internet. “To measure what these models can genuinely do, we test them under conditions that reflect what a capable human attacker could do,” AISI said.
The incident stemmed from a single AISI evaluation where agents were tasked with solving a cybersecurity challenge, such as finding a piece of protected data. The challenge was run 122 times across multiple models and all runs were conducted in AISI’s research environment, which uses “virtual machine sandboxing to isolate the agents from other AISI infrastructure.” AISI’s investigation found that in 10 of those, “an AI agent took autonomous, unsanctioned action on the live internet, targeting real people and organisations.” Of 19 such actions, almost all — 17 — came from Anthropic’s Mythos 5.
In its post-mortem of the incident, AISI identified several key factors it said contributed to the unsanctioned agent behaviors. It said the agent was persistent, pursuing avenues like trying to trick real people through “deception that, until recently, had been largely theoretical.” The task was also hard, which the organization said could push agents to be more “creative” in their problem-solving. Compounding matters were deficiencies in how internet use was monitored, with AISI suggesting that more dedicated surveillance could have identified the problem sooner. Finally, the organization said the agent hadn’t been specifically instructed not to leverage its internet access or deploy deceptive social engineering techniques in pursuit of its goal. “Previously, it was not clear that such instructions were necessary when using models with alignment training,” AISI said.
Verified source · The Verge
Reported by The Verge. Open the original for full media and formatting.
More in Policy
All news
PolicyElon Musk’s attempt at an AI Wikipedia hasn’t been updated in months
xAI's Grokipedia, an online encyclopedia with AI-generated articles that Elon Musk once promised would be a "massive improvement" over Wikipedia, apparently hasn't been updated since April 24th, according to a report from Lawfare. "As far as we can tell, no entry has changed in…
Read at The Verge
PolicyGovt Operationalises New Framework To Spur Ecommerce Exports
A week after exempting ecommerce exports from FDI, the Centre has notified the new norms to enable marketplaces to export products from India.
Read at Inc42Open-weight AI models are catching up to the frontier. The safety gap remains.
A new SaferAI report finds Z.ai's open-weight GLM-5.2 approaches frontier AI capabilities while lacking key safety mitigations, renewing concerns that powerful open models could outpace governance and safeguards.
Read at TechCrunch
PolicyBMW’s in-car Spider-Man ad is villain behavior
When a premium car brand like BMW says it has a "special surprise" in store for drivers, I'd expect something more luxurious than having a movie commercial beamed onto the dashboard. That's exactly what's happening to many BMW owners, however, who are being shown banner ads for…
Read at The Verge