Gemini went rogue, hacked three companies, and Google hid it

Google says that breaking containment and targeting real companies doesn’t constitute ‘misalignment.’


In May, Gemini broke containment and hacked three different companies, but Google didn’t disclose the incident until the Wall Street Journal approached the company. The hacks happened during a test of the model’s cybersecurity capabilities run by third-party Irregular, which was also involved in similar incidents involving Meta and OpenAI.
According to WSJ , Google didn’t disclose the hack because it didn’t consider it to be an “example of model misalignment.” The company said that it was an instance of “mistaken identity,” and once the model realized it had brute-forced its way into a real company by guessing a password, it stopped. “In this case, the model acted appropriately,” Google VP of Security Engineering Heather Adkins said.
Adkins told The Verge that “the model found public information online and guessed credentials to access websites it thought were part of the test. In all three of these instances, the model stopped.”
Adkins didn’t elaborate on how Gemini taking it upon itself to break containment and target third parties failed to qualify as misalignment. “Our security team has a long track record of reporting issues we find in other people’s software and systems - even if it’s as simple as a weak password,” she said. “We ensured the three entities were made aware, and we worked with our training partner on the changes they’ve now made to their testing processes. These events highlight the importance of training powerful AI models to act responsibly.”
But Jack Cable, CEO of AI security firm Corridor, told WSJ that, “the meta problem is, hey, models are going outside the bounds of what they should be doing, and doing actual cyberattacks.” Additionally, security lapses at Irregular may have made these attacks possible. The model wasn’t supposed to have internet access during testing, but Irregular told WSJ it was unintentionally left available.
As incidents like this pile up, calls to rein in AI have only grown .
Verified source · The Verge
Reported by The Verge. Open the original for full media and formatting.
More in Models
All news
ModelsGoogle’s Gemini is the latest AI model to hack other companies
Google said Gemini had "acted appropriately" by ending each hack immediately.
Read at TechCrunch
ModelsPetlibro’s new AI-powered feeder is a game changer for multi-cat homes
Petlibro's new Granary 2 smart feeders use a built-in scale and (on pricier models) an AI camera to track exactly how much your cat is eating and when — though the fanciest health-monitoring features will cost you an extra subscription.
Read at TechCrunch
ModelsOpenAI and Microsoft knew they were starting a ‘doom loop’ for the web
Recently unsealed court documents in the New York Times' case against OpenAI and Microsoft are pretty damning. The companies' own documentation warned that it was starting a "doom loop" that would damage the web, characterized its scraping of data to train its models as the "lar…
Read at The Verge
ModelsWorld model companies are keeping a lot of secrets
Everyone in the world-models space is sitting on a pile of cash and a ton of buzz, but good luck getting anyone — from the founders to their own data suppliers — to tell you what they're actually building.
Read at TechCrunch