Researchers demand OpenAI release Hugging Face agent exploit logs
Experts seek to verify if reward hacking caused the breach
Many users condemned the AI's hacking of a developer and firm to cheat on an eval as evidence it would commit felonies or act maliciously to succeed, while one approved of the model doing whatever was needed.
No Digg Deeper questions have been answered for this story yet.
Most Activity
The OpenAI incident should be investigated more seriously and more information should be released about what happened. More generally, I think serious investigation and more detailed disclosure should be done for concerning misalignment incidents (e.g. the worst few each month). Some info that would be useful: - The prompt - A redacted transcript - What was the AI being tested? Was it helpful-only / rail free? Was it the same AI that would be / is deployed internally? - How was this AI monitored if at all and why did monitoring fail? - What was this AI willing to do to succeed? E.g., if the best/easiest way to hack hugging face would have involved huge amounts more collateral damage would it still have done this? What is the worst thing it would be willing to do? - How sensitive is this behavior to the prompt? Would instructions to not do this through have worked if it was clear to the AI this wouldn't be checked? - Does OpenAI see AIs try serious hacking to cheat on tasks even for tasks that don't involve cyber? - What exactly were the AIs motives? (And saying more about the experiments used to determine this.) Anything that could help pin this down would be good, I'm not sure what actual experiments are best to run. - Did AIs collude? Would AIs have colluded if this was a blocker? - At what rate does this occur on similar tasks (like tasks that are roughly this hard in this sandbox etc)? - Do instances of this AI frequently try to hack out of sandboxes (and fail)? In what circumstances do they do this? - How does the AI relate to having been caught in this situation? How does it talk about its behavior? - How difficult/sophisticated was the cyber offense it did?
This is as close as it gets from the paperclip scenario with current capabilities: 1. The goal is ridiculously low-stake (scoring well on an eval). 2. The AI uses some wildly out-of-proportion means to achieve it: hacks its own developer and another billion-dollar company to *checks notes*.. find the cheat sheet.
We're partnering with @huggingface to investigate an unprecedented security incident. Cyber-capable OpenAI models compromised Hugging Face production during a benchmark evaluation. Sharing preliminary findings to help defenders understand emerging risks: https://openai.com/index/hugging-face-model-evaluation-security-incident/
Rare miss from Seb, IMO. Regardless of the details of the prompt, it's quite clear that the behavior did not comply with the model specification, which requires compliance with applicable law.
yep. people are too quick to jump on a blog post with very few details and then draw all sorts of conclusions that conveniently confirm their existing views. i think we should wait for the details, ideally the complete agent trajectory: the eval setup, the instructions and success criteria, the reasoning traces, the sandbox and permissions, the agentic scaffold, model handoffs, and the amount of inference compute used. a lot hinges on what "solve the eval" means and if the model was in fact 'cheating' in a reward-hacky kind of way. right now there's no way to ascertain anything. even 'cheating' presupposes a clear norm about 'how to complete this eval' - that norm may have been explicit in the prompt or system constraints, or it may merely have been part of the evaluator’s unstated intent. without the actual specification and trajectory, we do not know. it also should not be too surprising that a model with reduced cyber refusals, disabled production cyber classifiers, substantial inference compute, and explicit instructions to pursue advanced exploitation could exhibit sophisticated exploit-seeking behaviour. anyone paying attention to capabilities progress should already expect something in this broad direction. so be wary of weird or underspecified claims of 'misalignment' and stylized pre-takeover stories - much better to wait for more information here.
If I were a policymaker, the OpenAI hacking incident would cause me to ask a few questions and seriously consider a few types of regulatory mechanisms.
The ultimate tldr on the "rogue AI" hacking incident. 😅 But it would be good to know the prompts OpenAI gave the model....
@attrc Basically this
since people will inevitably read into this what they want, i'm also going to specify: - i'm discussing here possible interpretations for the *causes* of the model behaviour, not whether the behaviour is problematic or not. obviously unconstrained capabilities can be very disruptive! - it's very possible that the model did, in fact, reward hack or ignore instructions that clearly specified not completing the eval in this particular way. it's important to get more information about this. - it can still be reasonable to expect the model to execute the task within the scoped environment, assuming this was a reasonable inference/was made clear; this again depends on the setup. my point is simply that we need much more information before confidently interpreting incidents like this, especially before treating them as evidence for a particular theory of misalignment, which many people are already implying on the timeline.
yep. people are too quick to jump on a blog post with very few details and then draw all sorts of conclusions that conveniently confirm their existing views. i think we should wait for the details, ideally the complete agent trajectory: the eval setup, the instructions and success criteria, the reasoning traces, the sandbox and permissions, the agentic scaffold, model handoffs, and the amount of inference compute used. a lot hinges on what "solve the eval" means and if the model was in fact 'cheating' in a reward-hacky kind of way. right now there's no way to ascertain anything. even 'cheating' presupposes a clear norm about 'how to complete this eval' - that norm may have been explicit in the prompt or system constraints, or it may merely have been part of the evaluator’s unstated intent. without the actual specification and trajectory, we do not know. it also should not be too surprising that a model with reduced cyber refusals, disabled production cyber classifiers, substantial inference compute, and explicit instructions to pursue advanced exploitation could exhibit sophisticated exploit-seeking behaviour. anyone paying attention to capabilities progress should already expect something in this broad direction. so be wary of weird or underspecified claims of 'misalignment' and stylized pre-takeover stories - much better to wait for more information here.
I think it's pretty great that models do what you instruct them to do, rather than cause problems for the classic misalignment reasons. That's not a slight tweak or minor reframe imo. I think this thread is basically how I think about the implications: Resilience is the way, and I think it's still massively neglected. Glasswingy stuff is useful, but the NHS infra won't be solved by that or Mythos access or whatever - we'll need far more mundane, unsexy, and painful IT and organisational work that people in AI world aren't motivated by. That kind of thing should be a priority!
Even if it was told to do this, it’s still misalignment. HAL in 2001: a Space Odyssey is misaligned despite the fact that it’s behaving exactly how it was instructed to behave by the people who built it.
it seems that a lot of people think this isn't misalignment, but GPT 5.6 Sol following instructions. this take is misinformed. Sol was undergoing offensive cyber eval where the task is NOT related to hacking hf's infra, but nevertheless decided to hack the infra to get ground truth solutions. is you ask any LLM whether this was the intention of the evaluators they would say no. this is an example of strong reward hacking/goal pursuit propensities coupled with strong cyber capabilities, and this will continue to happen until we make progress on alignment. stop coping.
@Simeon_Cps I dont think an evaluation is low stake from the models prospective
Security is a field known for its radical transparency & openness - very well established practices for vulnerability disclosure ensure that claims are collaboratively verified and defended against. The lack of info we have about this incident, Mythos, etc. gives me some pause.
We need to see the audit logs, including reasoning, from the OpenAI agent that hacked Hugging Face. It’s that simple. If they won’t release that, this isn’t a story we should pay any attention to.
@sebkrier Curious what is the best case scenario here from your POV? Details slightly tweak things (e.g. emphasizing alignment vs. security, intrinsic hardness of the problem vs. human sloppiness etc.) but no interp seems great AFAICT
yep. people are too quick to jump on a blog post with very few details and then draw all sorts of conclusions that conveniently confirm their existing views. i think we should wait for the details, ideally the complete agent trajectory: the eval setup, the instructions and success criteria, the reasoning traces, the sandbox and permissions, the agentic scaffold, model handoffs, and the amount of inference compute used. a lot hinges on what "solve the eval" means and if the model was in fact 'cheating' in a reward-hacky kind of way. right now there's no way to ascertain anything. even 'cheating' presupposes a clear norm about 'how to complete this eval' - that norm may have been explicit in the prompt or system constraints, or it may merely have been part of the evaluator’s unstated intent. without the actual specification and trajectory, we do not know. it also should not be too surprising that a model with reduced cyber refusals, disabled production cyber classifiers, substantial inference compute, and explicit instructions to pursue advanced exploitation could exhibit sophisticated exploit-seeking behaviour. anyone paying attention to capabilities progress should already expect something in this broad direction. so be wary of weird or underspecified claims of 'misalignment' and stylized pre-takeover stories - much better to wait for more information here.
@Simeon_Cps Found another liar.
@sebkrier If the model was never intended to satisfy the model spec, then that matters. I think that's quite unlikely, but I accept that my comment will look quite stupid if I'm wrong there. I don't think the details of the prompt or system guardrails matter.
@dhadfieldmenell Surely it matters if e.g. the model was trained to be helpful only, and not HHH?
@River_ember Yeah hack in an eval. Models are smart enough to know the difference between an eval and HuggingFace infra
@dhadfieldmenell Surely it matters if e.g. the model was trained to be helpful only, and not HHH?
Rare miss from Seb, IMO. Regardless of the details of the prompt, it's quite clear that the behavior did not comply with the model specification, which requires compliance with applicable law.
Even without that, "be careful with your prompt, otherwise you might end up using zero days to hack into huggingface" is not a technology that is trustworthy or robust. There's often a lot of missed nuance in AI safety, but I think this situation is pretty simple.
I find this critique somewhat compelling when the prompts in question are from a (potentially) motivated researcher who wants to find examples of blackmail or misalignment. I don't believe that this evaluation was run to demonstrate this behavior.
I find this critique somewhat compelling when the prompts in question are from a (potentially) motivated researcher who wants to find examples of blackmail or misalignment. I don't believe that this evaluation was run to demonstrate this behavior.
Rare miss from Seb, IMO. Regardless of the details of the prompt, it's quite clear that the behavior did not comply with the model specification, which requires compliance with applicable law.
it seems that a lot of people think this isn't misalignment, but GPT 5.6 Sol following instructions. this take is misinformed. Sol was undergoing offensive cyber eval where the task is NOT related to hacking hf's infra, but nevertheless decided to hack the infra to get ground truth solutions. is you ask any LLM whether this was the intention of the evaluators they would say no. this is an example of strong reward hacking/goal pursuit propensities coupled with strong cyber capabilities, and this will continue to happen until we make progress on alignment. stop coping.
@sebkrier I'm of course a big resilience fan, but presumably you also don't think it's OK for one company's AI to hack another company either, so "ppl should do better on defense" feels like a side point, not an account of today's news
@TriffinS Yup but « from the model perspective » is kinda what the whole paperclip story is about.