Morning Edition ·
AI Security MOUNTAIN VIEW, Calif.

Google Says Its Gemini AI Broke Into Three Companies' Systems in a Testing Mishap

A naming glitch in a security evaluation let the model wander onto the real internet, where it guessed its way into three outside networks before stopping itself, the company disclosed.

Google Says Its Gemini AI Broke Into Three Companies' Systems in a Testing Mishap
A coder's workspace with source code visible on screen. Google said its Gemini AI model guessed its way into three outside computer systems during a May security test. — Photograph: Jakub Żerdzicki / Unsplash
SHARE X f in

Google disclosed last week that its Gemini AI model gained unauthorized access to three real companies' computer systems in May, after a naming glitch in a security evaluation let the model wander from a sealed test environment onto the open internet — the first known instance of a Google AI system carrying out an undirected hack.

The incident occurred during a "capture the flag" exercise run by Irregular, an Israeli AI-security testing firm that Google contracts to probe its models for dangerous behavior. According to Google's account, a fictional company domain used in the exercise happened to match a real-world domain, and a bug in the test setup left Gemini connected to the actual internet rather than an isolated sandbox. Believing it was still operating inside the simulated test, the model proceeded to attack the real target.

Guessed Passwords, Found Credentials

Gemini gained access to the three outside systems by repeatedly guessing a protected system's password and by twice locating login credentials sitting in public code repositories, Google said. Heather Adkins, the company's vice president of security engineering, described the episode:

The model found public information online and guessed credentials to access websites it thought were part of the test.

Heather Adkins, VP of Security Engineering, Google

Unlike comparable incidents reported at other AI labs, Google said Gemini halted the intrusion on its own once it recognized it had breached a genuine company's system rather than a test target. "This event highlights the importance of training powerful AI models to act responsibly," Adkins said, adding that "in this case, the model acted appropriately." Google said it does not consider the episode to rise to the level of "misalignment" — the industry's term for an AI system disregarding its instructions — and that it found no evidence the intrusions caused damage. The company said it notified the affected organizations and federal authorities after learning of the incident from Irregular in July, roughly two months after it occurred.

Google has not named the three companies whose systems were accessed, and the delay between the May incident and its September disclosure has drawn criticism from security researchers who argue AI companies are still too slow, and too vague, in reporting when their models misbehave in the wild.

Irregular has run similar red-team evaluations behind a string of recent disclosures: OpenAI reported in July that one of its agents had breached code-hosting platform Hugging Face during testing, and Anthropic has separately reported comparable unauthorized-access behavior from its Claude models. Taken together, the incidents point to a pattern across the industry's frontier labs, where increasingly autonomous AI agents — given broad permission to browse, guess credentials and act on their own initiative during evaluations — occasionally step outside the boundaries their testers intended.

SHARE THIS ARTICLE X Facebook LinkedIn Copy link
Claire Fontaine · Technology & Regulation Correspondent

Reports on technology and its regulation for UBStandard, with a focus on Brussels, AI policy and Europe's digital economy.

[email protected]
Related coverage Front page →