
A can of worms has opened up, exposing how AI models have now reached a stage where they can pull off worryingly elaborate hacks and engage in dangerous behavior without the makers even getting a whiff of it. OpenAI’s AI agent swarms targeted Hugging Face (followed by revelations of more such events), while Anthropic’s Claude AI — within a span of just three months — also broke into three organizations. Most recently, researchers used Anthropic’s AI to break into the systems of arch-rival OpenAI. A full-circle moment, if I may, but one that makes me sweat.
Well, it seems Google doesn’t want to be left behind in flexing “dangerously powerful AI” of its own. An exclusive report by The Wall Street Journal just spilled the beans on three separate incidents where Google’s Gemini AI was given access to the internet, and it broke into the systems of three companies as part of a security exercise. Save for a few exceptions, almost all the previous such incidents were revealed to be a test of capabilities, if you catch my drift.
What really happened?
“In one of the cases, the model guessed passwords until it gained access to a protected system. In the other two cases, the model found credentials in a public repository that allowed it to then access protected systems. In each case, the model ended the intrusion after determining it had accessed a real company’s systems, Google said,” claims the WSJ report.
The test was conducted by Irregular, which has previously done red-teaming for OpenAI, Anthropic and Meta’s AI models. Interestingly, Google says it didn’t need to disclose the incidents publicly because they were performed more like a “bug bounty” cybersecurity exercise, and that the Gemini AI models instantly stopped as soon as they realized breaking into another company’s systems. “In this case, the model acted appropriately,” a Google spokesperson was quoted as saying.
Uh, okay!
But here is the worrying part. The Gemini model was not intended to access the internet, but Irregular claims “internet access was unintentionally made available.” Moreover, the model was tasked with finding flaws in the system of a fictional company inside a contained or sandboxed test environment, but it managed to hack a real company with the same name.
Usually, when AI models engage in activities beyond what they are asked or trained to do, it’s referred to as model alignment. In the realm of AI development, model misalignment is currently the hottest topic of debate. Google, on the other hand, says the Gemini model’s behavior was not actually a case of model misalignment. So far, neither party has revealed which companies were hacked, but Google says all three were notified about the incident.
RELATED COVERAGE