If there's a public incident bad enough that it'd be pretty risky/hard for OpenAI *not* to disclose it, there were almost certainly more concerning incidents internally that we never heard about. Like, it'd be pretty surprising if the first time an AI hacks its way out of a sandbox and starts hacking other stuff, the thing it hacks is an external company rather than something inside OpenAI. First you get AIs hacking internal services (where disclosure isn't forced), and only later do you get something like this that's hard to keep quiet. So we probably could have seen this coming—if we'd known about the worst internal incidents. If there were a list of, say, the 10 worst incidents from a misalignment and severity perspective, you could look at how bad the worst one is and how fast severity falls off from there to get a real sense of how concerning things are. While OpenAI did disclose some earlier incidents, this was done in an ad hoc way such that we can't get a great sense of how bad things actually are. (And this was potentially too slow given how fast AI progress might go, though the delay is OK for now.) And of course, it's not clear that OpenAI is disclosing all risk-relevant details about this incident. We can't keep depending on ad hoc, voluntary disclosure as the stakes rise. And it seems pretty straightforward to do better: companies could maintain an updated list of the ~10 worst incidents over the past few months (with some delay—perhaps 1 or 2 weeks by default—before an incident has to be added and some allowed redactions). Better yet, a trusted third party could collect the worst incidents across all frontier companies, with a whistleblowing mechanism so employees can flag when the provided list/descriptions seriously misrepresent reality (and more generally avoid spin). A third party could also anonymize incidents—which also removes the incentive for companies to bury their heads in the sand.
Ryan Greenblatt argues public AI safety incidents follow worse internal events
Stella Biderman notes frontier labs already disclosed sandbox escapes.
Positive users agree with the proposal for third-party tracking of worst AI safety incidents, while negative users dismiss the idea by noting that sandbox escapes have already been disclosed by OpenAI and Anthropic.
No Digg Deeper questions have been answered for this story yet.
Most Activity
Note that I expect these concerns apply to all frontier AI companies, not just OpenAI:
@RyanGreenblatt Yeah, I kind of wonder how much of the public release was just a matter of getting ahead of a leak from the huggingface side, which had already reported information about the situation.
@RyanGreenblatt @dhadfieldmenell We know this isn’t the first time that an AI has hacked its way out of a sandbox because both OpenAI and Anthropic have previously disclosed that their models have hacked their way out of sandboxes during testing. Do you not read their model cards?
If there's a public incident bad enough that it'd be pretty risky/hard for OpenAI *not* to disclose it, there were almost certainly more concerning incidents internally that we never heard about. Like, it'd be pretty surprising if the first time an AI hacks its way out of a sandbox and starts hacking other stuff, the thing it hacks is an external company rather than something inside OpenAI. First you get AIs hacking internal services (where disclosure isn't forced), and only later do you get something like this that's hard to keep quiet. So we probably could have seen this coming—if we'd known about the worst internal incidents. If there were a list of, say, the 10 worst incidents from a misalignment and severity perspective, you could look at how bad the worst one is and how fast severity falls off from there to get a real sense of how concerning things are. While OpenAI did disclose some earlier incidents, this was done in an ad hoc way such that we can't get a great sense of how bad things actually are. (And this was potentially too slow given how fast AI progress might go, though the delay is OK for now.) And of course, it's not clear that OpenAI is disclosing all risk-relevant details about this incident. We can't keep depending on ad hoc, voluntary disclosure as the stakes rise. And it seems pretty straightforward to do better: companies could maintain an updated list of the ~10 worst incidents over the past few months (with some delay—perhaps 1 or 2 weeks by default—before an incident has to be added and some allowed redactions). Better yet, a trusted third party could collect the worst incidents across all frontier companies, with a whistleblowing mechanism so employees can flag when the provided list/descriptions seriously misrepresent reality (and more generally avoid spin). A third party could also anonymize incidents—which also removes the incentive for companies to bury their heads in the sand.
@RyanGreenblatt Agree with you here.
@RyanGreenblatt Well 2 days ago they did post about an incident from 2.5 months ago where it did hack out of the sandbox to post a PR to github. I think models just weren't capable enough until recently. Question is did anything else happen in these 2 months
@RyanGreenblatt What if their AI has hacked other companies/institutions and they're not owning up to it?
@BlancheMinerva @dhadfieldmenell "Starts hacking other stuff". Also prior examples I'm aware of are more mundane bypasses rather than finding zero-days.
This is a more detailed version of this earlier thread: