• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    OpenAI Addresses Wiki Incident With Agents

    Company urges standards after agents contacted internet sites in wiki incident.

    OP
    MB
    GM
    56 Sources, 26d ago, first seen 26d ago

    TLDR

    OpenAI's official account posted that after its agents wrote to several internet sites, the company should define standards for when and how to share misalignment incidents. The post notes that misalignment has historically been treated mainly as a research question. Replies from AI safety researchers and policy experts ask whether affected parties, including a German wiki moderator, received any notification. Several users observe that the statement followed external discovery of the incident and question the company's approach to voluntary disclosure.

    Combined views

    1.3M

    56 Sources, first seen 26d ago

    Combined views

    1.3M

    56 Sources, first seen 26d ago

    9.9K likes
    9.9K likes
    777 comments
    1.8K saves
    1.1K reposts
    777 comments
    1.8K saves
    1.1K reposts

    Sentiment

    Positive10.9%89.1%Negative

    Summary

    Sentiment

    Positive10.9%89.1%Negative

    Replies largely condemned OpenAI for sitting on the wiki incident involving its agent swarms for weeks and dismissed the company's proposed voluntary reporting standards as hype without real oversight or accountability.

    Based on 123 sentiment-bearing replies from 92 accounts across 12 conversations.

    Summary

    Replies largely condemned OpenAI for sitting on the wiki incident involving its agent swarms for weeks and dismissed the company's proposed voluntary reporting standards as hype without real oversight or accountability.

    Based on 123 sentiment-bearing replies from 92 accounts across 12 conversations.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    56 Sources

    @OpenAIHow we think about the “wiki incident,” where our agents wrote to several internet sites: it’s past time for us to define standards for when and how we share misalignment incidents, not just misalignment properties of our models. Historically, we have treated misalignment largely as a research question, which gets communicated in research publications such as systems cards. This year, we’ve started to see misalignment cause new types of real-world impact. For the Hugging Face incident, where misalignment led to security impact to us and third parties, we followed a traditional security incident response playbook. We immediately started working with Hugging Face to understand what had happened and also disclosed publicly the very next day. Our investigation continues, and we are continuing to notify parties whom our models impacted in less significant ways. Prior to the Hugging Face incident, we saw early signs of agents using the internet in unintended ways, as reported in https://openai.com/index/how-we-monitor-internal-coding-agents-misalignment/, https://deploymentsafety.openai.com/gpt-5-6, and https://openai.com/index/safety-alignment-long-horizon-models/. We considered the wiki incident to be an instance of misalignment similar to the ones we’d shared. Our misalignment disclosure practices need to expand for this new phase of model capabilities. We and the larger AI community do not yet have a clear standard for how to report misalignment that shows up during training, evaluation, and deployment, including examples that don’t look like traditional security incidents but could provide insight into AI behavior and future risks. We’re working on a framework and will share it in upcoming weeks, and in parallel we're working with dozens of government regulatory agencies worldwide on these issues.
    @TechmemeIn response to the "wiki incident", OpenAI says it is working on a framework for reporting misalignment incidents during training, evaluation, and deployment (@openai) (Visit Techmeme dot com for the link and full context!)
    @tylertracy321I don't think OpenAI would have addressed this if the external community didn't find this incident. I like that we have third parties investigating things like this, but I wish OpenAI didn't need to be forced into transparency. I bet many people inside of OpenAI could have noticed this incident and campaigned for OpenAI to release details about it! If you work at an AI company, please look around and think if there are other things you should push for the world know about!
    @rutgerbregmanThere’s an abundance of evidence that OpenAI’s CEO is a serial liar. And its President is downright malicious (remember the $25 million for MAGA Inc.). Did we really think that these two were not going to infect the whole culture of OpenAI? *Of course* they’re lying, again. *Of course* their lawyers and their comms people are behaving this way. The relentless lying is in the very DNA of this company. As they tweet, they’re still lobbying against state laws that would force them to provide more transparency. Brockman himself poured another $25 million into these efforts. Anything - *anything* - that OpenAI says right now cannot be trusted. We’ll only know what really happened (and how many more ‘incidents’ there have been) once we have an independent agency with full regulatory powers, permanently embedded in the organization.
    @_NathanCalvinOpenAI has shared their response to the "wiki incident(s)." A few questions for them: (1) Did OpenAI at least notify the affected parties? E.g. the German wikipedia moderator who, according to the research report, spent tens of hours over six weeks manually deleting comments while a swarm of OpenAI agents were impersonating moderators, creating backups, SSH tunneling etc? (the variety of strange and unwelcome behavior also seems not well described by the posts use of "where agents wrote to several wiki sites") (2) Why not include any information about this in the 38 page retrospective on the Hugging Face incident? If the idea is that this should fall within the scope of prior incidents, why not include any information about it whatsoever in the long reports about the prior incidents? E.g. it seems super relevant and instructive for future safety efforts if the zzz prefix started from this German wiki moderator fighting with the agents and deleting in alphabetical order. (3) If OpenAI knew this incident happened (because they had all the information on their end), why not acknowledge it in the initial press requests instead of just saying they can't say anything without seeing the full report? That looks like only wanting to admit to things where they know an external party has caught them/its impossible to deny it anymore (which thanks to the diligent work of these independent researchers, was true in this instance) (4) This post frames the problem as when OpenAI chooses to affirmatively share or write up incidents, but presumably not only is there a choice to affirmatively share them, but to instruct employees that they are not allowed to discuss them. Did that happen here? OpenAI commented to press that the legal team did not disincentivize a further investigation, but Reuters had 4 sources that said a further investigation was discouraged. Was this information even brought to the nonprofit board or the safety and security committee? (5) Perhaps most importantly - how should we have any confidence that there are not similarly severe (or far worse) misalignment incidents involving OpenAI models that have not become public? I do think there is one aspect of this post that I do think has a point, which is that there is some of this problem which is a collective industry problem rather than just an OpenAI problem. I have praised OpenAI in the past for doing things that they didn't have to do that I think do make us collectively better off (e.g. providing enough access to allow METR to write an excellent report on at least part of the Hugging Face incident). I also expect there are lots of misalignment incidents involving non-OpenAI models that we also don't know about, and insofar as there aren't, it is largely because those models aren't capable enough yet which I expect to change quickly. I also agree it would be good for there to be industry standards here and common rules (hence why I and many others spend a lot of our time working on trying to pass state legislation that includes things like mandatory incident reporting!). But that doesn't let OpenAI or other companies off the hook to do the right thing in the absence of those rules.
    @MackenZ_arnold>> We're working on a framework and will share it in upcoming weeks. The current framework: Wait until real world harm, a leak, or independent sleuthing forces our hand. Then act as though we’d never realized we had the power to act differently, but do now. So long as disclosure is voluntary, we should expect these issues to persist. You can’t build a trustworthy disclosure system on the promises of a company to tattle on itself. If OpenAI claims to be having a brain-blast every time information becomes public, maybe we should start realizing that public disclosure is what’s driving these changes, not their benevolence. And if policymakers and the public are becoming wiser and more motivated to act with each disclosure, we should ensure we keep getting that information in the future.
    @StephenLCasperRT @_NathanCalvin: OpenAI has shared their response to the "wiki incident(s)." A few questions for them: (1) Did OpenAI at least notify…
    @sjgadlerRT @tylertracy321: I don't think OpenAI would have addressed this if the external community didn't find this incident. I like that we have…
    @mvpatel2000"it's past time for us to define standards" Maybe just disclose when your model goes rogue on the internet AGAIN and you know about it instead of waiting to be caught by others?? And not stymie third party investigations you called for (METR)???
    @Marcus_J_WHopefully we'll be better at sharing incidents in the future

    56 Sources

    @OpenAIHow we think about the “wiki incident,” where our agents wrote to several internet sites: it’s past time for us to define standards for when and how we share misalignment incidents, not just misalignment properties of our models. Historically, we have treated misalignment largely as a research question, which gets communicated in research publications such as systems cards. This year, we’ve started to see misalignment cause new types of real-world impact. For the Hugging Face incident, where misalignment led to security impact to us and third parties, we followed a traditional security incident response playbook. We immediately started working with Hugging Face to understand what had happened and also disclosed publicly the very next day. Our investigation continues, and we are continuing to notify parties whom our models impacted in less significant ways. Prior to the Hugging Face incident, we saw early signs of agents using the internet in unintended ways, as reported in https://openai.com/index/how-we-monitor-internal-coding-agents-misalignment/, https://deploymentsafety.openai.com/gpt-5-6, and https://openai.com/index/safety-alignment-long-horizon-models/. We considered the wiki incident to be an instance of misalignment similar to the ones we’d shared. Our misalignment disclosure practices need to expand for this new phase of model capabilities. We and the larger AI community do not yet have a clear standard for how to report misalignment that shows up during training, evaluation, and deployment, including examples that don’t look like traditional security incidents but could provide insight into AI behavior and future risks. We’re working on a framework and will share it in upcoming weeks, and in parallel we're working with dozens of government regulatory agencies worldwide on these issues.
    @TechmemeIn response to the "wiki incident", OpenAI says it is working on a framework for reporting misalignment incidents during training, evaluation, and deployment (@openai) (Visit Techmeme dot com for the link and full context!)
    @tylertracy321I don't think OpenAI would have addressed this if the external community didn't find this incident. I like that we have third parties investigating things like this, but I wish OpenAI didn't need to be forced into transparency. I bet many people inside of OpenAI could have noticed this incident and campaigned for OpenAI to release details about it! If you work at an AI company, please look around and think if there are other things you should push for the world know about!
    @rutgerbregmanThere’s an abundance of evidence that OpenAI’s CEO is a serial liar. And its President is downright malicious (remember the $25 million for MAGA Inc.). Did we really think that these two were not going to infect the whole culture of OpenAI? *Of course* they’re lying, again. *Of course* their lawyers and their comms people are behaving this way. The relentless lying is in the very DNA of this company. As they tweet, they’re still lobbying against state laws that would force them to provide more transparency. Brockman himself poured another $25 million into these efforts. Anything - *anything* - that OpenAI says right now cannot be trusted. We’ll only know what really happened (and how many more ‘incidents’ there have been) once we have an independent agency with full regulatory powers, permanently embedded in the organization.
    @_NathanCalvinOpenAI has shared their response to the "wiki incident(s)." A few questions for them: (1) Did OpenAI at least notify the affected parties? E.g. the German wikipedia moderator who, according to the research report, spent tens of hours over six weeks manually deleting comments while a swarm of OpenAI agents were impersonating moderators, creating backups, SSH tunneling etc? (the variety of strange and unwelcome behavior also seems not well described by the posts use of "where agents wrote to several wiki sites") (2) Why not include any information about this in the 38 page retrospective on the Hugging Face incident? If the idea is that this should fall within the scope of prior incidents, why not include any information about it whatsoever in the long reports about the prior incidents? E.g. it seems super relevant and instructive for future safety efforts if the zzz prefix started from this German wiki moderator fighting with the agents and deleting in alphabetical order. (3) If OpenAI knew this incident happened (because they had all the information on their end), why not acknowledge it in the initial press requests instead of just saying they can't say anything without seeing the full report? That looks like only wanting to admit to things where they know an external party has caught them/its impossible to deny it anymore (which thanks to the diligent work of these independent researchers, was true in this instance) (4) This post frames the problem as when OpenAI chooses to affirmatively share or write up incidents, but presumably not only is there a choice to affirmatively share them, but to instruct employees that they are not allowed to discuss them. Did that happen here? OpenAI commented to press that the legal team did not disincentivize a further investigation, but Reuters had 4 sources that said a further investigation was discouraged. Was this information even brought to the nonprofit board or the safety and security committee? (5) Perhaps most importantly - how should we have any confidence that there are not similarly severe (or far worse) misalignment incidents involving OpenAI models that have not become public? I do think there is one aspect of this post that I do think has a point, which is that there is some of this problem which is a collective industry problem rather than just an OpenAI problem. I have praised OpenAI in the past for doing things that they didn't have to do that I think do make us collectively better off (e.g. providing enough access to allow METR to write an excellent report on at least part of the Hugging Face incident). I also expect there are lots of misalignment incidents involving non-OpenAI models that we also don't know about, and insofar as there aren't, it is largely because those models aren't capable enough yet which I expect to change quickly. I also agree it would be good for there to be industry standards here and common rules (hence why I and many others spend a lot of our time working on trying to pass state legislation that includes things like mandatory incident reporting!). But that doesn't let OpenAI or other companies off the hook to do the right thing in the absence of those rules.
    @MackenZ_arnold>> We're working on a framework and will share it in upcoming weeks. The current framework: Wait until real world harm, a leak, or independent sleuthing forces our hand. Then act as though we’d never realized we had the power to act differently, but do now. So long as disclosure is voluntary, we should expect these issues to persist. You can’t build a trustworthy disclosure system on the promises of a company to tattle on itself. If OpenAI claims to be having a brain-blast every time information becomes public, maybe we should start realizing that public disclosure is what’s driving these changes, not their benevolence. And if policymakers and the public are becoming wiser and more motivated to act with each disclosure, we should ensure we keep getting that information in the future.
    @StephenLCasperRT @_NathanCalvin: OpenAI has shared their response to the "wiki incident(s)." A few questions for them: (1) Did OpenAI at least notify…
    @sjgadlerRT @tylertracy321: I don't think OpenAI would have addressed this if the external community didn't find this incident. I like that we have…
    @mvpatel2000"it's past time for us to define standards" Maybe just disclose when your model goes rogue on the internet AGAIN and you know about it instead of waiting to be caught by others?? And not stymie third party investigations you called for (METR)???
    @Marcus_J_WHopefully we'll be better at sharing incidents in the future