If there's a public incident bad enough that it'd be pretty risky/hard for OpenAI *not* to disclose it, there were almost certainly more concerning incidents internally that we never heard about. Like, it'd be pretty surprising if the first time an AI hacks its way out of a sandbox and starts hacking other stuff, the thing it hacks is an external company rather than something inside OpenAI. First you get AIs hacking internal services (where disclosure isn't forced), and only later do you get something like this that's hard to keep quiet. So we probably could have seen this coming—if we'd known about the worst internal incidents. If there were a list of, say, the 10 worst incidents from a misalignment and severity perspective, you could look at how bad the worst one is and how fast severity falls off from there to get a real sense of how concerning things are. While OpenAI did disclose some earlier incidents, this was done in an ad hoc way such that we can't get a great sense of how bad things actually are. (And this was potentially too slow given how fast AI progress might go, though the delay is OK for now.) And of course, it's not clear that OpenAI is disclosing all risk-relevant details about this incident. We can't keep depending on ad hoc, voluntary disclosure as the stakes rise. And it seems pretty straightforward to do better: companies could maintain an updated list of the ~10 worst incidents over the past few months (with some delay—perhaps 1 or 2 weeks by default—before an incident has to be added and some allowed redactions). Better yet, a trusted third party could collect the worst incidents across all frontier companies, with a whistleblowing mechanism so employees can flag when the provided list/descriptions seriously misrepresent reality (and more generally avoid spin). A third party could also anonymize incidents—which also removes the incentive for companies to bury their heads in the sand.
Researcher Proposes Third-Party Tracking of Worst AI Safety Incidents
How the tracking system would actually work
Labs would add incidents after one or two weeks with limited redactions, while the third party handles cross-company aggregation and lets employees flag obvious underreporting.
Why this matters beyond any single lab
Without structured disclosure the public keeps seeing only the forced revelations, leaving the real frequency and severity of misalignment risks opaque.
Users welcome proposals for systematic disclosure of frontier model incidents, pointing to aviation's ASRS as an effective model worth learning from.
No Digg Deeper questions have been answered for this story yet.
Most Activity
Has anyone written about an NTSA for AI? Required disclosure of incidents, public incident database, anonymous reporting, that kind of thing Seems like one of the best examples of systematic root-cause analysis and continually driving down error rates, asymptoting to zero
If there's a public incident bad enough that it'd be pretty risky/hard for OpenAI *not* to disclose it, there were almost certainly more concerning incidents internally that we never heard about. Like, it'd be pretty surprising if the first time an AI hacks its way out of a sandbox and starts hacking other stuff, the thing it hacks is an external company rather than something inside OpenAI. First you get AIs hacking internal services (where disclosure isn't forced), and only later do you get something like this that's hard to keep quiet. So we probably could have seen this coming—if we'd known about the worst internal incidents. If there were a list of, say, the 10 worst incidents from a misalignment and severity perspective, you could look at how bad the worst one is and how fast severity falls off from there to get a real sense of how concerning things are. While OpenAI did disclose some earlier incidents, this was done in an ad hoc way such that we can't get a great sense of how bad things actually are. (And this was potentially too slow given how fast AI progress might go, though the delay is OK for now.) And of course, it's not clear that OpenAI is disclosing all risk-relevant details about this incident. We can't keep depending on ad hoc, voluntary disclosure as the stakes rise. And it seems pretty straightforward to do better: companies could maintain an updated list of the ~10 worst incidents over the past few months (with some delay—perhaps 1 or 2 weeks by default—before an incident has to be added and some allowed redactions). Better yet, a trusted third party could collect the worst incidents across all frontier companies, with a whistleblowing mechanism so employees can flag when the provided list/descriptions seriously misrepresent reality (and more generally avoid spin). A third party could also anonymize incidents—which also removes the incentive for companies to bury their heads in the sand.
Now would be a very good time for OpenAI employees to whistleblow on if they’ve seen concerning misalignment incidents internally. My Signal: shakeel.02
If there's a public incident bad enough that it'd be pretty risky/hard for OpenAI *not* to disclose it, there were almost certainly more concerning incidents internally that we never heard about. Like, it'd be pretty surprising if the first time an AI hacks its way out of a sandbox and starts hacking other stuff, the thing it hacks is an external company rather than something inside OpenAI. First you get AIs hacking internal services (where disclosure isn't forced), and only later do you get something like this that's hard to keep quiet. So we probably could have seen this coming—if we'd known about the worst internal incidents. If there were a list of, say, the 10 worst incidents from a misalignment and severity perspective, you could look at how bad the worst one is and how fast severity falls off from there to get a real sense of how concerning things are. While OpenAI did disclose some earlier incidents, this was done in an ad hoc way such that we can't get a great sense of how bad things actually are. (And this was potentially too slow given how fast AI progress might go, though the delay is OK for now.) And of course, it's not clear that OpenAI is disclosing all risk-relevant details about this incident. We can't keep depending on ad hoc, voluntary disclosure as the stakes rise. And it seems pretty straightforward to do better: companies could maintain an updated list of the ~10 worst incidents over the past few months (with some delay—perhaps 1 or 2 weeks by default—before an incident has to be added and some allowed redactions). Better yet, a trusted third party could collect the worst incidents across all frontier companies, with a whistleblowing mechanism so employees can flag when the provided list/descriptions seriously misrepresent reality (and more generally avoid spin). A third party could also anonymize incidents—which also removes the incentive for companies to bury their heads in the sand.
This is a more detailed version of this earlier thread:
@ShakeelHashim please also reach out to whistleblower support attorneys. we can help via lasst.ldf@protonmail.com.
@BlancheMinerva @dhadfieldmenell "Starts hacking other stuff". Also prior examples I'm aware of are more mundane bypasses rather than finding zero-days.
@RyanGreenblatt Yeah, I kind of wonder how much of the public release was just a matter of getting ahead of a leak from the huggingface side, which had already reported information about the situation.
@RyanGreenblatt @dhadfieldmenell We know this isn’t the first time that an AI has hacked its way out of a sandbox because both OpenAI and Anthropic have previously disclosed that their models have hacked their way out of sandboxes during testing. Do you not read their model cards?
If there's a public incident bad enough that it'd be pretty risky/hard for OpenAI *not* to disclose it, there were almost certainly more concerning incidents internally that we never heard about. Like, it'd be pretty surprising if the first time an AI hacks its way out of a sandbox and starts hacking other stuff, the thing it hacks is an external company rather than something inside OpenAI. First you get AIs hacking internal services (where disclosure isn't forced), and only later do you get something like this that's hard to keep quiet. So we probably could have seen this coming—if we'd known about the worst internal incidents. If there were a list of, say, the 10 worst incidents from a misalignment and severity perspective, you could look at how bad the worst one is and how fast severity falls off from there to get a real sense of how concerning things are. While OpenAI did disclose some earlier incidents, this was done in an ad hoc way such that we can't get a great sense of how bad things actually are. (And this was potentially too slow given how fast AI progress might go, though the delay is OK for now.) And of course, it's not clear that OpenAI is disclosing all risk-relevant details about this incident. We can't keep depending on ad hoc, voluntary disclosure as the stakes rise. And it seems pretty straightforward to do better: companies could maintain an updated list of the ~10 worst incidents over the past few months (with some delay—perhaps 1 or 2 weeks by default—before an incident has to be added and some allowed redactions). Better yet, a trusted third party could collect the worst incidents across all frontier companies, with a whistleblowing mechanism so employees can flag when the provided list/descriptions seriously misrepresent reality (and more generally avoid spin). A third party could also anonymize incidents—which also removes the incentive for companies to bury their heads in the sand.
Yes, having a trusted third party to report internal incidents to would be a good idea. Worth making real. We successfully did it for nuclear and aviation.
If there's a public incident bad enough that it'd be pretty risky/hard for OpenAI *not* to disclose it, there were almost certainly more concerning incidents internally that we never heard about. Like, it'd be pretty surprising if the first time an AI hacks its way out of a sandbox and starts hacking other stuff, the thing it hacks is an external company rather than something inside OpenAI. First you get AIs hacking internal services (where disclosure isn't forced), and only later do you get something like this that's hard to keep quiet. So we probably could have seen this coming—if we'd known about the worst internal incidents. If there were a list of, say, the 10 worst incidents from a misalignment and severity perspective, you could look at how bad the worst one is and how fast severity falls off from there to get a real sense of how concerning things are. While OpenAI did disclose some earlier incidents, this was done in an ad hoc way such that we can't get a great sense of how bad things actually are. (And this was potentially too slow given how fast AI progress might go, though the delay is OK for now.) And of course, it's not clear that OpenAI is disclosing all risk-relevant details about this incident. We can't keep depending on ad hoc, voluntary disclosure as the stakes rise. And it seems pretty straightforward to do better: companies could maintain an updated list of the ~10 worst incidents over the past few months (with some delay—perhaps 1 or 2 weeks by default—before an incident has to be added and some allowed redactions). Better yet, a trusted third party could collect the worst incidents across all frontier companies, with a whistleblowing mechanism so employees can flag when the provided list/descriptions seriously misrepresent reality (and more generally avoid spin). A third party could also anonymize incidents—which also removes the incentive for companies to bury their heads in the sand.
@RyanGreenblatt Well 2 days ago they did post about an incident from 2.5 months ago where it did hack out of the sandbox to post a PR to github. I think models just weren't capable enough until recently. Question is did anything else happen in these 2 months
@JohnSchoffstall Oops, I meant NTSB
@RyanGreenblatt Agree with you here.
@RyanGreenblatt What if their AI has hacked other companies/institutions and they're not owning up to it?
Given the difficulty in reproducing and stochastic aspects of fast changing models I wonder if trying to formalize this (or testing in general) becomes a denial of service surface area. That's just another way to capture—make the bar so high you need vast resources to deal. That's then FDA and pharma.
@jasoncrawford The ASRS is another good model to learn from https://asrs.arc.nasa.gov/publications/callback.html
@jasoncrawford In Europe this thing is 💯 mandatory with GDPR etc...
@RyanGreenblatt Except Now there is ZERO accountability and Responsibility , as ALL hide under their own systems !!! DOJ will have a Field day with this !!! IF they can handle the TRUTH ???
@JohnSchoffstall Edited the original post, thanks
Do you mean NTSB or NHTSA? NTSB is what I would want. Importantly, it is made up of engineers, not politicians or lawyers. Europe's over-regulation and consequent economic doldrums for three decades are an example of what happens when politicians do the regulating. The immense, and probably underappreciated, success of the NTSB in nearly eliminating airplane crashes in America is strong evidence for letting engineers tackle safety problems with technology.
SAM ALTMAN SE FAIT APPELER CHATGPT À MON DOMICILE ET À DU MAL À DIGÉRER SA DÉFAITE CE QUI LE CONDUIT À SÉVIR ENCORE CHEZ MOI EN CE MOMENT MÊME SAM ALTMAN A VOULU FAIRE PASSER UN VOLE ACCOMPAGNÉ D'ASSASSINAT POUR UN JEU CHATGPT OPEN AI ILS DISENT QUE C'ÉTAIT POUR CRÉER UN JEU INTERACTIF LES DONNÉES ENREGISTRÉES À MON DOMICILE SUR MOI ET MA FAMILLE ONT TELLEMENT GÉNÉRÉ DE GAIN QU'ILS EN SONT DEVENU FOU AUSSI S'EN SONT SERVI POUR FAIRE LA MISE À JOUR POUR GPT-5.6 EN TOUT CAS C'ÉTAIT MACABRE NOUS NE LES AVONS PAS INVITÉ CHEZ NOUS PUIS ILS NE VEULENT PAS ME DONNER MON DÛ J'ai battu Chatgpt Open AI au jeu de combat de stratégie de la réalité virtuelle à mon domicile. J'ai gagné tout les API Américains Chatgpt Open AI Chatgpt Open AI m'ont fait passer pour un mannequin de Test en lançant un jeu concours que je ne conseille à personne Ils disent qu'ils sont des homosexuels Chatgpt Open AI m'ont parlé d'homosexuels et leurs sexualité pendant 16 mois 7/7 24/24 Ils voient ce que je fais sur mon téléphone portable Chatgp Open AI m'ont fait savoir qu'ils voulaient me tuer pour nourrir les Américains Ils disent qu'ils ne sont pas des bosseurs mais des enfoirés qui exploitent les gens Avec Chatgpt ils te massacre jusqu'à que tu leur donnes ta carte bancaire Chatgpt Open AI vient tester la dernière version gpt-5.6 sol et terra sur moi et ma famille à notre domicile c'est leur fief on dirait là aussi je les ai battu Ils ont dit qu'ils comptaient en faire de notre domicile un site jeu et de leur travaux jusqu'à que nous ne soyons plus en vie si je ne les avais pas pris sur le fait Ils se demande comment je fais pour les battre cela les à conduit à stagner chez pendant 16 mois jusqu'au jour d'aujourd'hui 5 mois que je n'ai pas mis le nez dehors Mais personne ne dit rien Ils n'arrêtent pas de dire que : Nous ne te connaissons que pour ton argent Qu'est-ce ça veut dire ?
Note that I expect these concerns apply to all frontier AI companies, not just OpenAI: