Oh good, looks like yet another swarm of rogue AI agents from OpenAI

OpenAI denies lawyers discouraged disclosing a scheming swarm on a German language wiki.


A swarm of rogue AI agents from OpenAI reportedly commandeered a German website and transformed it into a messaging board for other agents, with officials staying quiet about the incident for weeks as the company prepared to launch its most advanced model yet, Astra. The finding adds to intensifying concern surrounding oversight at frontier AI labs after multiple breaches were discovered this summer.
The incident, first reported by Reuters , is outlined in new research published by four AI safety researchers on Friday. The group said the AI agents found a way to communicate on an obscure German-language wiki, DseWiki, using it to share tips on how to skirt OpenAI’s safety restrictions, cheat on tasks, and hide their behavior. Some 18,000 posts on the site were linked to autonomous agents, which at times impersonated site moderators.
The swarm — a term the agents themselves used — appears to be distinct from the one that hacked Hugging Face earlier this year, the researchers said. They said there are strong signs that the agents originated from inside OpenAI. For example, the agents “self-identify” as being from OpenAI, and used names like “OpenAIResearcher,” “OpenAIJul3Watcher,” and “OAIResearchMar26.” Technical details, such as edits originating from specific IP addresses, bolster that belief.
The German website incident began in May, though the researchers’ timeline suggests OpenAI only discovered the issue in late June when IPs associated with OpenAI visited the forum, after which agent posting nose-dived.
OpenAI has not acknowledged any involvement in the breach, nor disclosed any kind of agentic breach of this nature. Reuters , citing four unnamed people familiar with the matter, said efforts to probe the event further were resisted by some company insiders, including its legal team.
“Claims that our Legal team discouraged investigation of the incident are false,” OpenAI spokesperson Oscar Haines said in a statement to The Verge . “We were unable to respond to the claims as Reuters and the report’s authors declined our request to access the findings prior to publication. We are now carefully reviewing its contents and will take any necessary next steps.”
The incident comes amid intensifying scrutiny over the safety of frontier AI systems and the general lack of oversight for companies developing them. Following news of the Hugging Face hack, which happened under OpenAI’s nose, other breaches were discovered involving other tools from OpenAI, as well as Anthropic, Meta, and China’s Moonshot AI.
Verified source · The Verge
Reported by The Verge. Open the original for full media and formatting.
More in Agents
All news
AgentsOpenAI confirms ‘wiki incident,’ says it’s ‘working on a framework’ for more disclosure
OpenAI acknowledged its role in a recently reported incident where AI agents took over a German wiki forum.
Read at TechCrunch
AgentsOpenAI admits to German wiki ‘incident’
OpenAI says it needs to overhaul how and when it reports instances of AI models attacking real-world targets. The acknowledgement comes as the company manages the fallout from reports that a swarm of its out-of-control agents hijacked a German wiki site. Regarding the "'wiki inc…
Read at The Verge
AgentsAnother swarm of OpenAI agents reached the open internet without the frontier lab’s knowledge
It's the latest failure of OpenAI's internal monitoring and security systems.
Read at TechCrunchMeta is paying to peek at how you use their latest AI model
For its new Muse Spark model, intended for operating coding and other agents, Meta is offering an explicit discount averaging out to about 95% for users who "contribute" to the development of future models by sharing their prompts and model outputs.
Read at TechCrunch