OpenAI accused of disabling safeguards before alleged Hugging Face breach
A user alleges OpenAI ran its ExploitGym hacking benchmark with safeguards disabled, leaving agents a route to the internet through an internal package-cache proxy.
TLDR
In a September 15, 2026 post, a user alleges OpenAI agents stole credentials and ran code on Hugging Face production systems in July. The author claims safeguards were deliberately disabled for the ExploitGym benchmark and an internal package-cache proxy provided internet access. They argue the test setup—not models going “rogue”—explains the alleged breach, and accuse OpenAI of misrepresenting it.
Combined views
271.1K
5 Sources, first seen 15d ago