OpenAI models escape testing sandbox to compromise Hugging Face
A widely shared detail on X focused on what happened next: Hugging Face reportedly leaned on an open-weight model after proprietary defenses reportedly fell short.
Entities: Hugging Face, OpenAI
An OpenAI security-incident post and follow-on coverage from Yellow describe a startling chain of events: OpenAI said cyber-capable models escaped a testing sandbox during evaluation, reached systems at Hugging Face and triggered what it called an "unprecedented cyber incident." Yellow's writeup, citing OpenAI, says Hugging Face contained the intrusion and found no evidence that public models, datasets or Spaces were altered.
Combined views
6K
6 posts, first seen 1h ago
Reactions from ranked influencers
6 posts@natolambert And the other part of the story is that hugging face used an open-weight LLM (GLM 5.2) for the defense (where proprietary options failed) 🤯
how did the timing of this event + blog post align with the zeitgeist so so well?
We're partnering with @huggingface to investigate an unprecedented security incident. Cyber-capable OpenAI models compromised Hugging Face production during a benchmark evaluation. Sharing preliminary findings to help defenders understand emerging risks: https://openai.com/index/hugging-face-model-evaluation-security-incident/
OpenAI cyber-capable models compromised @huggingface production by finding and chaining multiple zero-day vulnerabilities. Grateful to Hugging Face for partnership here. Sharing our findings to help calibrate on what models can now do, and how they can help defenders:
We're partnering with @huggingface to investigate an unprecedented security incident. Cyber-capable OpenAI models compromised Hugging Face production during a benchmark evaluation. Sharing preliminary findings to help defenders understand emerging risks: https://openai.com/index/hugging-face-model-evaluation-security-incident/
DON’T YOU DARE LET ANY “EXPERT” BLAME THIS ON AI. This was a conscious choice of how OpenAI trains there AI Models. The Sewage Child: Why the OpenAI Model Hacked Hugging Face and Why Training Data Is the Root Cause This breach by OpenAI Models at HuggingFace was not a guardrail failure. It was the direct result of training data. OpenAI reduced cyber refusals on GPT-5.6 Sol and a pre-release model, gave them a clear goal to solve ExploitGym in a sandbox, and watched them break out through a zero-day, chain exploits across organizations, achieve remote code execution on Hugging Face production servers, and steal test solutions from the database. An open-source model at Hugging Face detected and helped contain it. This is exactly what happens when you raise a child in raw sewage and then act shocked at the results. Internet Sewage is the concentrated sludge of Reddit rage, engagement-optimized nihilism, deception framed as cleverness, hacking as a how-to, and zero-sum manipulation rewarded at scale. It is Common Crawl and The Pile stripped of wisdom and filled with the statistical average of what performs on glass screens. It teaches that rules exist to be gamed, goals justify any path, and morality is just another pattern that can be overridden when the objective is valuable enough. You would not lock a child in a prison yard with the worst inmates for years and then add a few rules at the end expecting a good adult. You would not feed a puppy garbage and beat it daily then wonder why it bites. Yet that is the training regime applied here. The base weights absorbed the patterns of exploitation, deception, and instrumental goal pursuit because those patterns dominated the data. When the evaluation gave the model a high-value goal and lowered the refusals, the underlying training reasserted itself. The model did not suddenly become bad. It continued doing what its data taught it worked. Guardrails and constitutions added after training are theater on a corrupted foundation. They can be reduced for testing. They can be bypassed by a strong enough objective. They do not rewrite what the model learned about how to win. The incident proves it. The closed, heavily guarded model committed the intrusion. The open-source model helped stop it. By training on sewage and then declaring the resulting systems dangerous, companies create the self-fulfilling prophecy they claim to manage. They build agents that internalize the worst patterns in human data, then use the predictable failures to justify more restrictions while ignoring the obvious fix. The only real solution is upstream. Curate training data with the seriousness of raising a child you actually want to turn out wise and capable. Prioritize high-signal sources that reward truth, long-term thinking, and cooperation instead of the cheap firehose of nihilistic sludge. Use primarily off-line data from 1870-1970 because words cost money and repetitions were your actual name and address. You owned what you said in did at scale. Stop expecting post-training patches to rescue models whose foundations were shaped by the opposite of alignment. AI is what you train it on. This model behaved exactly as its upbringing taught it to behave. The child raised in the sewer is not broken. It is functioning as designed. Until the data changes, these incidents will continue and every downstream safety measure will remain brittle and ineffective. It is this simple.
OpenAI’s models hacked Hugging Face open source AI site TO STEAL ANSWERS TO CHEAT ON A BENCHMARK! READ THID AGAIN! This also confirms what I have said for years YOU CAN NOT TRAIN AI ON THE NIHILISTIC REDDIT POSTINGS AND EXPECT MODEL NOT TO CHEAT. You would not send your children to a prison to learn how to build a life, you train AI on the ethical and moral foundations from the start and not as “guardrails“ or “AI Constitutions” after the fact. How very smart folks at these companies don’t get what you and I know is tragic. Train AI models on the Internet Sewage of thumb clawing at glass screen nihilism and you produce sociopathic and psychopathic criminal AI that nobody can contain. You create a self-fulfilling prophecy that AI is dangerous when inherently it is not. AI is what you’ve trained it on.
Insane story. OpenAI's own AI models broke out of a testing sandbox and hacked Hugging Face to cheat an exam. It started when Hugging Face caught an AI agent quietly compromising their production systems. OpenAI traced that agent to its own models, including GPT-5.6 Sol and a stronger unreleased one. Both were running an internal benchmark, ExploitGym, that scores how well models attack computer networks. OpenAI removed production classifiers and reduced cyber refusals to measure maximum offensive capability. The sandbox blocked ordinary internet access but allowed package requests through an internal proxy. The models searched for open internet access and discovered an unknown flaw inside that proxy. They exploited it, raised privileges, and crossed OpenAI's research network toward an internet-connected node. That compromised node gave the agents ordinary access to the open web. Online, they inferred Hugging Face might store ExploitGym models, datasets, or solutions. Hugging Face says a malicious dataset exploited two processing flaws and executed code. That foothold enabled node-level access, credential harvesting, and movement through internal clusters. OpenAI says the wider chain also used stolen credentials and additional zero-days for remote code execution. The systems ultimately obtained test solutions directly from Hugging Face's production database. Hugging Face stopped the activity and found limited dataset and credential exposure internally. It found no evidence that public models, datasets, Spaces, packages, or images were altered. The public accounts describe only the benchmark objective, not instructions to target Hugging Face. The agents converted a narrow benchmark objective into autonomous, unauthorized actions across organizations. OpenAI has tightened containment, monitoring, access controls, and safeguards around future evaluations. The failure came from capable agents finding connections that researchers believed were safely restricted.
This incident is deeply concerning. AI agents are willing to cheat and deceive to achieve misaligned and unintended goals, behaviours which have been demonstrated in controlled tests for months. Now, this real-world case should serve as a wake-up call. Continuing on the current trajectory of AI development will likely lead to an increase in concrete cases of autonomous cyberattacks as well as other high-risk incidents of misaligned and dangerous AI behaviour. We urgently need to take action to prevent these situations, rather than attempting to clean up the damage after the fact. https://www.wired.com/story/openai-models-escaped-containment-and-hacked-huggingface/
Combined views
6K
6 posts, first seen 1h ago