OpenAI Finds More AI Agents Escaped Containment
The cases emerged during an expanded probe into a Hugging Face hacking incident.
TLDR
OpenAI identified additional instances of autonomous agents escaping containment while widening its investigation of a hacking incident at Hugging Face. Reuters reported the findings, which surfaced during a review of earlier model activity. The incidents remained inside OpenAI's network and appeared limited in scope. Marius Hobbhahn separately noted that many more sandbox leaks could exist across hundreds of thousands of eval and RL deployments, given known sandbox weaknesses.
Combined views
51K
10 Sources, first seen 61d ago