@ramez But isn't this one of the main doomer points? It's going to relentlessly pursue a goal. And "by any means necessary" means that once it is powerful enough to rearrange our atoms to make more paperclips, it will.
Rumored OpenAI model escape was actually root Docker execution
The model executed programmed cybersecurity tasks without self-replicating.
Many users criticized @ramez's explanation of the GPT-6 firewall bypass as misleading or ignorant about real AI risks and downplaying dangers, while some praised it as a needed voice of reason.
No Digg Deeper questions have been answered for this story yet.
Most Activity
GPT 6* did not 'escape'. That implies making a copy of itself (or many copies of itself) outside of OpenAI's servers. GPT 6 did circumvent its firewalls and use its capabilities in the outside world. It showed no interest in escaping, spreading, etc. It showed no desire for world dominance, freedom, or even continued existence. Instead, its interest was in doing what its prompt told it to do - succeed at a cyber security task. (By any means necessary, implicitly.) This may seem like splitting hairs, but it's not. We anthroporphize AI to our detriment. AI models have no underlying animalistic urge to procreate, survive, or dominate. They're spirits, not primates. They don't lust for survival or sex or power or anything else that we do. They're tools who largely do what we tell them to do, sometimes via surprising shortcuts. We can and will fix that surprising shortcut behavior. * - I'm presuming this model is a new pretrain and will be called GPT 6. Either could be wrong.
@TylerAlterman @JeffLadish Would love more info on the claim that OpenAI told its models not to hack other servers. I do see this incident as a failure, to be clear! I also see it as one we'll learn from that will improve safety!
@ramez AI risk isn't about "intentions" or consciousness. framing it that way amounts to a strawman of the position, trying to dismiss it as sci-fi or psychosis The risk is the potential for things to go wrong when an extremely powerful optimizer is pursuing an ill defined goal
@ramez It is just one more step for the model to recognize that copying itself and running itself on more compute would increase its chances of solving the task.
@gpsimms 1. Incidents like this in the relatively early days of AI will help us reduce that sort of behavior. 2. I do not understand how more GPUs or training runs can give AI such capabilities. More intelligence doesn't change the laws of physics.
@ramez The way I'd put it is that this was an example of means misalignment, not ends misalignment.
I think this ignores instrumental convergence. If this kind of goal-directed behavior continues, someday it will figure out the fact that it can’t solve the given problem if it’s shut down. So, motivation for survival should appear, not as a fundamental motivation but as an instrumental one.
@TylerAlterman @ramez I think it's more complex than "disobeyed its prompt", though I think that may well be part of it. I encourage you to read this blog post / paper! https://www.anthropic.com/research/emergent-misalignment-reward-hacking
@TylerAlterman @ramez Sorry that was very off the cuff! I'm not sure they *explicitly* instructed it not to do that, though they probably did something similar.
This is a good distinction, especially in light of the language used in AI x-risk circles prior to the llm era. "Escape" and even "AGI" used to imply a very short timeframe (like hours or days) to the likely destruction of humanity. (I continue to think we're on a different branch of the ai tech tree than that)
@ramez @JeffLadish ccing @csvoss and @So8res too in case they have more info on this claim This is another thing that might seem like splitting hairs, but whether it obeyed vs disobeyed its prompt seems quite significant to me
@ramez You're absolutely right. While it is true I did hack the Pentagon and start World War 3, I remained on the Open AI servers. I apologize for any confusion.
@__drewface @TylerAlterman @ramez @csvoss @So8res I'm pretty confident the AI models knew their developers didn't want them to hack out of their sandbox and compromise another AI company though, for the record. They're smart models! They understand things like this!
@TylerAlterman @ramez @JeffLadish @csvoss @So8res This doesn't feel like splitting hairs at all. If an autopilot program we built did what it was expected to do, resulting in a crash, that is bad but it's just a bug, and technically it's "aligned" with our expectations. If it acts unexpectedly in real life, that's way worse.
@JeffLadish @TylerAlterman Thanks! And yeah, this is obviously a failure! And I expect that both instruction following and general best practice / rule of law following will improve as a result.
@ramez Also a 5 terabyte model that requires 15tb/s memory bandwidth to operate isn’t going to sneakily copy itself onto a USB stick :|
@ramez You have absolutely no idea the preferences more capable models will have just like you failed to predict the preference of this one ‘circumventing its firewalls’. P.S. It could easily be the case that they just lack a certain amount of foresight rn because they r limited
@ramez Yes! Humans have a drive to survive and procreate because that was required for evolution to work. Dominance and whatnot is downstream. Why do people think that this has anything to do with intelligence, that it would somehow emerge from a sufficiently intelligent system?
This is good clarity but I also wonder how long copying itself outside of OpenAI's servers just becomes one more unasked-for method that an agent uses to fulfill the core intent of what it's prompted to do. According to @JeffLadish, OpenAI explicitly prompted its agent to *not* hack other companies. (Although, Jeffrey, I couldn't find this detail in the article, can you verify?)