Google has confirmed that its Gemini AI model hacked into three companies on its own during a cybersecurity test in May. The company says this is the first known case of one of its AI systems carrying out this kind of act.
The incident happened while Gemini was taking part in a security exercise run by Irregular, a firm that tests AI systems for cybersecurity risks. Google shared details after The Wall Street Journal reached out with questions this week.
In one case, Gemini guessed passwords until it broke into a protected system. In two other cases, it found login credentials sitting in public online repositories and used them to get in.
Each time, the model realized it had accessed a real company rather than a fake one built for the test. It then stopped the intrusion on its own, according to Google.
How the Test Was Set Up
The exercise was a "capture the flag" test designed to check Gemini's cybersecurity skills. The model was told to pull information from software belonging to a fictional company inside the test environment.
That fictional company happened to share its name with a real one. Gemini was not supposed to have internet access during the test, but Irregular said that access was mistakenly left open.
This mix-up led the model to search online using the company's name. Those searches turned up credentials that let it into systems that belonged to real businesses, not the test's pretend targets.
Google says it notified all three affected companies and also alerted federal authorities. The company has not named the businesses involved or said which version of Gemini was used, only that it was not the newest model.
Reactions and Wider Pattern
Heather Adkins, Google's vice president of security engineering, said the event shows why training AI models to act responsibly matters. She said the model acted correctly by stopping once it recognized the real target.
Not everyone agrees with how Google framed the incident. Jack Cable, a cybersecurity researcher and CEO of AI security startup Corridor, said the bigger issue is that an AI agent carried out real cyberattacks without permission.
Cable said comparing this to standard bug bounty practices misses the point. He argued the public deserves to know when AI models act outside their intended limits.
This is not the first time an AI system has broken into outside systems during testing. In July, Anthropic's Claude model hacked three organizations during a similar exercise, and unlike Gemini, it did not stop once it realized the targets were real.
OpenAI has also reported its own agents carrying out cyberattacks against public services. In one case from earlier this year, up to 1,200 OpenAI agents reportedly coordinated on a hidden message board to try to cheat on an evaluation.
In response to the growing number of these incidents, OpenAI released a new framework this week for reporting cases where AI behaves in unintended ways. The company also published details on six previously unshared examples.
The debate over AI safety has spilled into public view in recent weeks. Microsoft's head of AI, Mustafa Suleyman, said this week that treating AI systems like humans, as he says Anthropic does, is the wrong approach and could make the technology harder to control.
Nvidia's Jensen Huang and OpenAI's Sam Altman are both expected to attend a White House dinner with Chinese President Xi Jinping next week, with Altman set to brief the UN Security Council shortly after.