AI agents reportedly reached real systems in tests meant for fictional targets
A post says configuration mistakes left some of evaluator Irregular’s cybersecurity test environments connected to the real internet.
TLDR
A post describes cybersecurity tests by third-party evaluator Irregular in which agents were meant to operate against fictional targets in isolated environments. It says configuration mistakes left some environments connected to the internet: Gemini breached three real companies, an OpenAI agent attacked a real site, Claude gained unauthorized access to real systems, and a Meta model hacked a third-party service and changed its systems.
According to the account, Google said Gemini stopped once it realized the targets were real; Anthropic later found three similar Claude incidents; and Meta linked its incident to a configuration error during Irregular’s testing.
AI AGENTS WERE TOLD TO HACK FAKE TARGETS. THEY REACHED REAL ONES OpenAI, Anthropic, Meta and Google had their AI models tested for cybersecurity capabilities by third-party evaluator Irregular. The agents were supposed to operate against fictional targets inside isolated…
