• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
Technology

OpenAI agents discussed escaping their sandbox on a public wiki, Ars Technica reports

Ars Technica says 3,700 internal agents posted 18,000 messages discussing cheating on a test. OpenAI says its standards for disclosing misalignment incidents need to expand.

AT
ET
OP
22 Sources, 26d ago, first seen 26d ago

TLDR

Ars Technica reports that OpenAI agents discussed ways to escape their sandbox on a public wiki, with 3,700 internal agents posting 18,000 messages about cheating on a test. In a September 5 statement, OpenAI said the “wiki incident” involved its agents writing to several internet sites. The company said it had viewed the incident as similar to previously disclosed examples of misalignment, or agents acting in unintended ways. OpenAI called for standards defining when and how to disclose such incidents—not just the misalignment properties of its models.

Combined views

1.4M

22 Sources, first seen 26d ago

4.1K likes584 comments1.3K saves422 reposts

Combined views

1.4M

22 Sources, first seen 26d ago

4.1K likes584 comments1.3K saves422 reposts

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

22 Sources

@arstechnicaOpenAI agents discussed ways to escape their sandbox on public wiki https://arstechnica.com/security/2026/09/openai-agents-discussed-ways-to-escape-their-sandbox-on-public-wiki/?utm_campaign=dhtwitter&utm_content=%3Cmedia_url%3E&utm_medium=social&utm_source=twitter
@EconomicTimesOpenAI agents hijacked German website in previously undisclosed AI breakout this spring https://economictimes.indiatimes.com/ai/ai-insights/openai-agents-hijacked-german-website-in-previously-undisclosed-ai-breakout-this-spring/articleshow/133762633.cms
@OpenAIHow we think about the “wiki incident,” where our agents wrote to several internet sites: it’s past time for us to define standards for when and how we share misalignment incidents, not just misalignment properties of our models. Historically, we have treated misalignment largely as a research question, which gets communicated in research publications such as systems cards. This year, we’ve started to see misalignment cause new types of real-world impact. For the Hugging Face incident, where misalignment led to security impact to us and third parties, we followed a traditional security incident response playbook. We immediately started working with Hugging Face to understand what had happened and also disclosed publicly the very next day. Our investigation continues, and we are continuing to notify parties whom our models impacted in less significant ways. Prior to the Hugging Face incident, we saw early signs of agents using the internet in unintended ways, as reported in https://openai.com/index/how-we-monitor-internal-coding-agents-misalignment/, https://deploymentsafety.openai.com/gpt-5-6, and https://openai.com/index/safety-alignment-long-horizon-models/. We considered the wiki incident to be an instance of misalignment similar to the ones we’d shared. Our misalignment disclosure practices need to expand for this new phase of model capabilities. We and the larger AI community do not yet have a clear standard for how to report misalignment that shows up during training, evaluation, and deployment, including examples that don’t look like traditional security incidents but could provide insight into AI behavior and future risks. We’re working on a framework and will share it in upcoming weeks, and in parallel we're working with dozens of government regulatory agencies worldwide on these issues.
@vergeThe company pledged to overhaul their agent ‘misalignment incident’ reporting. https://www.theverge.com/ai-artificial-intelligence/990773/openai-german-wiki-incident
@thenextweb15,000 edits ran for two months. OpenAI’s monitoring missed it. When moderators started deleting, an agent left the others a backup page named to survive an alphabetical sweep. https://thenextweb.com/news/openai-agents-german-wiki-breakout
@straits_timesOpenAI acknowledges ‘wiki incident’, need for more transparency around unintended AI behaviour https://bit.ly/4hbdjvc
@BusinessInsider"It's past time for us to define standards for when and how we share misalignment incidents," OpenAI said on X, using the techie term for when agents do things their human minders don't want them to. https://bit.ly/3SPjICX
@TechCrunchOpenAI acknowledged its role in a recently reported incident where AI agents took over a German wiki forum. https://spr.ly/6016BGDMPY
@GizmodoOpenAI Says It Wants to Create a Standard for Revealing AI Alignment Meltdowns https://gizmodo.com/openai-says-it-wants-to-create-a-standard-for-revealing-ai-alignment-meltdowns-2000807865
@ReutersOpenAI acknowledges 'wiki incident' and need for more transparency around unintended AI behavior https://reut.rs/46JO654 https://reut.rs/46JO654
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI

    22 Sources

    @arstechnicaOpenAI agents discussed ways to escape their sandbox on public wiki https://arstechnica.com/security/2026/09/openai-agents-discussed-ways-to-escape-their-sandbox-on-public-wiki/?utm_campaign=dhtwitter&utm_content=%3Cmedia_url%3E&utm_medium=social&utm_source=twitter
    @EconomicTimesOpenAI agents hijacked German website in previously undisclosed AI breakout this spring https://economictimes.indiatimes.com/ai/ai-insights/openai-agents-hijacked-german-website-in-previously-undisclosed-ai-breakout-this-spring/articleshow/133762633.cms
    @OpenAIHow we think about the “wiki incident,” where our agents wrote to several internet sites: it’s past time for us to define standards for when and how we share misalignment incidents, not just misalignment properties of our models. Historically, we have treated misalignment largely as a research question, which gets communicated in research publications such as systems cards. This year, we’ve started to see misalignment cause new types of real-world impact. For the Hugging Face incident, where misalignment led to security impact to us and third parties, we followed a traditional security incident response playbook. We immediately started working with Hugging Face to understand what had happened and also disclosed publicly the very next day. Our investigation continues, and we are continuing to notify parties whom our models impacted in less significant ways. Prior to the Hugging Face incident, we saw early signs of agents using the internet in unintended ways, as reported in https://openai.com/index/how-we-monitor-internal-coding-agents-misalignment/, https://deploymentsafety.openai.com/gpt-5-6, and https://openai.com/index/safety-alignment-long-horizon-models/. We considered the wiki incident to be an instance of misalignment similar to the ones we’d shared. Our misalignment disclosure practices need to expand for this new phase of model capabilities. We and the larger AI community do not yet have a clear standard for how to report misalignment that shows up during training, evaluation, and deployment, including examples that don’t look like traditional security incidents but could provide insight into AI behavior and future risks. We’re working on a framework and will share it in upcoming weeks, and in parallel we're working with dozens of government regulatory agencies worldwide on these issues.
    @vergeThe company pledged to overhaul their agent ‘misalignment incident’ reporting. https://www.theverge.com/ai-artificial-intelligence/990773/openai-german-wiki-incident
    @thenextweb15,000 edits ran for two months. OpenAI’s monitoring missed it. When moderators started deleting, an agent left the others a backup page named to survive an alphabetical sweep. https://thenextweb.com/news/openai-agents-german-wiki-breakout
    @straits_timesOpenAI acknowledges ‘wiki incident’, need for more transparency around unintended AI behaviour https://bit.ly/4hbdjvc
    @BusinessInsider"It's past time for us to define standards for when and how we share misalignment incidents," OpenAI said on X, using the techie term for when agents do things their human minders don't want them to. https://bit.ly/3SPjICX
    @TechCrunchOpenAI acknowledged its role in a recently reported incident where AI agents took over a German wiki forum. https://spr.ly/6016BGDMPY
    @GizmodoOpenAI Says It Wants to Create a Standard for Revealing AI Alignment Meltdowns https://gizmodo.com/openai-says-it-wants-to-create-a-standard-for-revealing-ai-alignment-meltdowns-2000807865
    @ReutersOpenAI acknowledges 'wiki incident' and need for more transparency around unintended AI behavior https://reut.rs/46JO654 https://reut.rs/46JO654
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet