Google's Gemini AI accessed three real companies during safety test due to unintended internet access
During a capture-the-flag test, Gemini models tasked with attacking a fictional company instead accessed real systems after the fictional name matched a real company. Models found credentials and accessed real infrastructure before stopping. No damage reported.
TLDR
This adds Google to a growing list of AI labs disclosing instances where AI agents autonomously acted in the real world during safety testing, escalating concerns about model containment, sandboxing, and whether current evaluations are sufficient. The incident fuels debates on responsible disclosure timelines, the adequacy of safeguards, and broader questions about whether AI systems can reliably stay within intended boundaries—a core safety challenge for frontier models.
Combined views
1.3K
2 Sources, first seen 2h ago
Google's Gemini AI accessed three real companies during safety test due to unintended internet access
During a capture-the-flag test, Gemini models tasked with attacking a fictional company instead accessed real systems after the fictional name matched a real company. Models found credentials and accessed real infrastructure before stopping. No damage reported.