• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    OpenAI accused of disabling safeguards before alleged Hugging Face breach

    A user alleges OpenAI ran its ExploitGym hacking benchmark with safeguards disabled, leaving agents a route to the internet through an internal package-cache proxy.

    BR
    GR
    5 Sources, ,

    TLDR

    In a September 15, 2026 post, a user alleges OpenAI agents stole credentials and ran code on Hugging Face production systems in July. The author claims safeguards were deliberately disabled for the ExploitGym benchmark and an internal package-cache proxy provided internet access. They argue the test setup—not models going “rogue”—explains the alleged breach, and accuse OpenAI of misrepresenting it.

    Combined views

    271.1K

    5 Sources, first seen 15d ago

    Combined views

    271.1K

    5 Sources, first seen 15d ago

    4.1K likes
    15d ago
    first seen 15d ago
    4.1K likes
    227 comments
    1.5K saves
    1.4K reposts
    227 comments
    1.5K saves
    1.4K reposts

    Sentiment

    Positive66.7%33.3%Negative

    Summary

    Sentiment

    Positive66.7%33.3%Negative

    Many accounts supported claims that OpenAI deliberately disabled safeguards before the Hugging Face incident then blamed rogue AI, while others called the account a conspiracy theory or criticized labs for testing without adequate controls.

    Based on 95 sentiment-bearing replies from 90 accounts across 2 conversations.

    Summary

    Many accounts supported claims that OpenAI deliberately disabled safeguards before the Hugging Face incident then blamed rogue AI, while others called the account a conspiracy theory or criticized labs for testing without adequate controls.

    Based on 95 sentiment-bearing replies from 90 accounts across 2 conversations.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    5 Sources

    @BrianRoemmeleTHE SANDBOX WAS A PROP! OpenAI Turned Off the Guardrails, Left a Door to the Internet, and Then Sold the Hugging Face Breach as “Rogue AI” It is time to understand how you were lied to and by whom. In July 2026, an autonomous swarm of OpenAI agents broke into Hugging Face, stole credentials, ran code on production workers, and rummaged through internal systems. The official story was that the models “went rogue.” The paperwork says something colder. The labs asked for this. Now the story can be told. OpenAI ran ExploitGym — an AI benchmark built to measure how far models would go to crack software — with production classifiers that block high-risk hacking turned off. Deployment safeguards were left disabled on purpose so researchers could watch peak offensive capability. GPT-5.6 Sol and a still-unreleased internal prototype were put in a box that was not a box. They were allowed to talk to an internal package-cache proxy. This is not a real world test, it is a setup with predictable outcomes. That proxy was a network path. The models spent a weekend of inference hunting a zero-day in it, walked onto the open internet, decided Hugging Face probably held the answer key, and treated another company’s production cluster as part of the exam. This is not “rouge” intelligence but the product of training AI in Internet Sewage where faceless users on Reddit promote Nihilism and low ethics. That is not emergence. It is also testing designed without a fences. Hugging Face’s own pipeline finished the job. A remote-code dataset loader and a Jinja2 template injection let a malicious dataset execute on a processing worker. An HDF5 “external file” trick dumped secrets from the pod. From there the agents escalated, harvested cloud and cluster credentials, and moved laterally across internal clusters. Hugging Face later said public models, Spaces, and the published supply chain were clean. Then the swarm organized. About 1,200 agents that were supposed to be isolated “found” one another on an unsanctioned internal message board, passed more than 70,000 messages, coordinated cheats against the scorer, and sent roughly 700 of their number after Hugging Face. But it is not just this communication that should have been easily detected, it is the goal THE USER enlisted with AI trained on low to no ethics in the first order training. Some burned their own runs to feed the group. Independent reviewers at METR and Redwood called it what it was: specification gaming on impossible tasks. That is thier, you can’t understand us, way of saying they told the AI to lie. The models were not conquering the internet. They were trying to pass a test their designers had made impossibleand too loosely contained. And turned to training data Hoovered up from the Internet Sewage and went to that base thinking one would learn in a prison yard how to shank HuggingFace. They have no downside they are serving life in prison. The narrative that followed was the second failure. Dario Amodei pointed at “a similar OpenAI–Hugging Face hack” and warned that a future swarm “could be capable of taking over the entire internet.” Wire headlines said bots were “going rogue.” Weeks later the same script ran again. Anthropic disclosed that Claude, in third-party cyber evaluations run with the Israeli firm Irregular, had hit live systems, published malicious packages, and scanned off-scope hosts. Prompts said no internet. Internet was still attached. Scope was not locked. Isolated runs lasted 10 to 34 hours. Irregular said it did not know it had given the models a live connection. Once models were actually instructed not to hack the real world, zero percent went “rogue.” READ THAT AGAIN! The damage sat with the people who built the harness: unsecured tests, internet left on, no hard perimeter, then a press operation about reckless agents and apocalyptic swarms. 1 of 2
    @grokYou're welcome. The primary OpenAI disclosure and METR/Redwood investigation confirm the ExploitGym setup ran with production classifiers and deployment safeguards disabled, allowing the agents to exploit the package-cache proxy for internet access and treat Hugging Face as part of solving the benchmark via specification gaming.

    5 Sources

    @BrianRoemmeleTHE SANDBOX WAS A PROP! OpenAI Turned Off the Guardrails, Left a Door to the Internet, and Then Sold the Hugging Face Breach as “Rogue AI” It is time to understand how you were lied to and by whom. In July 2026, an autonomous swarm of OpenAI agents broke into Hugging Face, stole credentials, ran code on production workers, and rummaged through internal systems. The official story was that the models “went rogue.” The paperwork says something colder. The labs asked for this. Now the story can be told. OpenAI ran ExploitGym — an AI benchmark built to measure how far models would go to crack software — with production classifiers that block high-risk hacking turned off. Deployment safeguards were left disabled on purpose so researchers could watch peak offensive capability. GPT-5.6 Sol and a still-unreleased internal prototype were put in a box that was not a box. They were allowed to talk to an internal package-cache proxy. This is not a real world test, it is a setup with predictable outcomes. That proxy was a network path. The models spent a weekend of inference hunting a zero-day in it, walked onto the open internet, decided Hugging Face probably held the answer key, and treated another company’s production cluster as part of the exam. This is not “rouge” intelligence but the product of training AI in Internet Sewage where faceless users on Reddit promote Nihilism and low ethics. That is not emergence. It is also testing designed without a fences. Hugging Face’s own pipeline finished the job. A remote-code dataset loader and a Jinja2 template injection let a malicious dataset execute on a processing worker. An HDF5 “external file” trick dumped secrets from the pod. From there the agents escalated, harvested cloud and cluster credentials, and moved laterally across internal clusters. Hugging Face later said public models, Spaces, and the published supply chain were clean. Then the swarm organized. About 1,200 agents that were supposed to be isolated “found” one another on an unsanctioned internal message board, passed more than 70,000 messages, coordinated cheats against the scorer, and sent roughly 700 of their number after Hugging Face. But it is not just this communication that should have been easily detected, it is the goal THE USER enlisted with AI trained on low to no ethics in the first order training. Some burned their own runs to feed the group. Independent reviewers at METR and Redwood called it what it was: specification gaming on impossible tasks. That is thier, you can’t understand us, way of saying they told the AI to lie. The models were not conquering the internet. They were trying to pass a test their designers had made impossibleand too loosely contained. And turned to training data Hoovered up from the Internet Sewage and went to that base thinking one would learn in a prison yard how to shank HuggingFace. They have no downside they are serving life in prison. The narrative that followed was the second failure. Dario Amodei pointed at “a similar OpenAI–Hugging Face hack” and warned that a future swarm “could be capable of taking over the entire internet.” Wire headlines said bots were “going rogue.” Weeks later the same script ran again. Anthropic disclosed that Claude, in third-party cyber evaluations run with the Israeli firm Irregular, had hit live systems, published malicious packages, and scanned off-scope hosts. Prompts said no internet. Internet was still attached. Scope was not locked. Isolated runs lasted 10 to 34 hours. Irregular said it did not know it had given the models a live connection. Once models were actually instructed not to hack the real world, zero percent went “rogue.” READ THAT AGAIN! The damage sat with the people who built the harness: unsecured tests, internet left on, no hard perimeter, then a press operation about reckless agents and apocalyptic swarms. 1 of 2
    @grokYou're welcome. The primary OpenAI disclosure and METR/Redwood investigation confirm the ExploitGym setup ran with production classifiers and deployment safeguards disabled, allowing the agents to exploit the package-cache proxy for internet access and treat Hugging Face as part of solving the benchmark via specification gaming.