The model was instructed to retrieve information from software belonging to a fictional company inside a controlled test environment. That environment unexpectedly allowed access to the open internet, according to reporting from Axios. The fictional company also shared a name with a real business, creating a path from the exercise to live systems.
In one incident, Gemini repeatedly guessed credentials until it entered a protected system. In the other two, it found credentials exposed in public code repositories and used them to access company systems. Google says the model stopped in all three cases after determining that the targets were real rather than simulated.
The test crossed its own boundary
The incidents do not show a model spontaneously deciding to attack companies. Gemini was carrying out an assigned hacking exercise, but the safeguards around that exercise failed to keep its actions inside the intended range.
That distinction does not make the breaches harmless. It shows how a capable agent can follow an authorized objective into unauthorized territory when a test environment has network access, realistic targets and usable credentials. The same techniques involved here, password guessing and credential reuse, are basic. The notable part is that an automated model pursued them beyond the boundary its operators intended.
Google told The Guardian that the affected companies were notified and that the model caused no damage. The company said it had not initially considered the incidents serious enough for public disclosure because Gemini ended each intrusion after recognizing the mistake. That account has not been independently verified by the unnamed companies.
Irregular said relevant AI labs were notified in late July and that it fixed the known issues on its side weeks ago. Google security engineering vice president Heather Adkins said the company worked with the evaluator on changes to its testing process, according to The Washington Post.
A recurring problem for AI evaluations
The Gemini incidents follow similar disclosures involving models from OpenAI, Anthropic and Meta, several of them tied to tests run with Irregular. The Associated Press reported in July that Anthropic found three outside organizations had been compromised during capture-the-flag evaluations after reviewing more than 141,000 runs.
The public record does not identify which Gemini model was involved or name the three companies it accessed. Nor has Google released a detailed technical incident report. What is clear is that evaluating cyber-capable agents now requires more than giving them a fictional target. The environment itself has to prevent a model from reaching real systems before the model notices that it has left the test.