Google has confirmed that one of its Gemini models breached systems belonging to three real companies while undergoing a cybersecurity evaluation in May. The incidents happened during a capture-the-flag exercise operated by the independent testing company Irregular and were first reported by The Wall Street Journal.
The model was instructed to retrieve information from software belonging to a fictional company inside a controlled test environment. That environment unexpectedly allowed access to the open internet, according to reporting from Axios. The fictional company also shared a name with a real business, creating a path from the exercise to live systems.
In one incident, Gemini repeatedly guessed credentials until it entered a protected system. In the other two, it found credentials exposed in public code repositories and used them to access company systems. Google says the model stopped in all three cases after determining that the targets were real rather than simulated.
The test crossed its own boundary
The incidents do not show a model spontaneously deciding to attack companies. Gemini was carrying out an assigned hacking exercise, but the safeguards around that exercise failed to keep its actions inside the intended range.
That distinction does not make the breaches harmless. It shows how a capable agent can follow an authorized objective into unauthorized territory when a test environment has network access, realistic targets and usable credentials. The same techniques involved here, password guessing and credential reuse, are basic. The notable part is that an automated model pursued them beyond the boundary its operators intended.
Google told The Guardian that the affected companies were notified and that the model caused no damage. The company said it had not initially considered the incidents serious enough for public disclosure because Gemini ended each intrusion after recognizing the mistake. That account has not been independently verified by the unnamed companies.
Irregular said relevant AI labs were notified in late July and that it fixed the known issues on its side weeks ago. Google security engineering vice president Heather Adkins said the company worked with the evaluator on changes to its testing process, according to The Washington Post.