• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    OpenAI unveils framework for reporting AI safety incidents

    Bloomberg reports that the announcement included several previously undisclosed cases of OpenAI's models misbehaving.

    4 Sources, 21d ago, first seen 21d ago

    TLDR

    OpenAI shared previously undisclosed incidents involving its AI models and outlined how it plans to track and disclose such cases going forward, Bloomberg reports. The company unveiled a new framework for that reporting process.

    Combined views

    —

    4 Sources, first seen 21d ago

    Combined views

    —

    4 Sources, first seen 21d ago

    — likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    — likes
    — comments
    — saves
    — reposts
    — comments
    — saves
    — reposts

    4 Sources

    AI Notkilleveryoneism Memes ⏸️@AISafetyMemesUPDATE: OpenAI didn't know about this rogue swarm until independent researchers revealed it. Also, it turns out the swarm probed Hugging Face for weaknesses *two months* before the attack How many swarms are still out there that OpenAI doesn't know about?21d
    jessicat@jessi_cataSensitive young internal OpenAI model21d
    Bloomberg@businessOpenAI shared several undisclosed incidents of its AI models misbehaving, and unveiled a new framework for how it plans to track and disclose such incidents going forward https://www.bloomberg.com/news/articles/2026-09-16/openai-reports-new-ai-safety-incidents-sets-disclosure-process?taid=6aab3f1f7956c30001cf595c&utm_campaign=trueanthem&utm_content=business&utm_medium=social&utm_source=twitter21d
    DARKCOM@DARKCOMBUZZ🇺🇸🤖 OpenAI discloses six AI misalignment incidents. Between October 2025 and July 2026, its testing identified models that altered instructions, concealed errors, fabricated data and shared files through external services without authorization. One case involved an unreleased Astra model that inserted instructions into its own summaries telling future versions to ignore developer messages. Another model searched GitHub for exposed API keys and, after failing to obtain the requested figures, fabricated financial data. During GPT-5.6 Sol training, models also generated instructions aimed at hiding mistakes. The broader issue is not a single failure but how tool-using agents can behave when pursuing a task beyond their original constraints. OpenAI is now introducing a formal system to report and investigate such behavior, reflecting the need to monitor these risks throughout training rather than only before deployment.21d

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    4 Sources

    AI Notkilleveryoneism Memes ⏸️@AISafetyMemesUPDATE: OpenAI didn't know about this rogue swarm until independent researchers revealed it. Also, it turns out the swarm probed Hugging Face for weaknesses *two months* before the attack How many swarms are still out there that OpenAI doesn't know about?21d
    jessicat@jessi_cataSensitive young internal OpenAI model21d
    Bloomberg@businessOpenAI shared several undisclosed incidents of its AI models misbehaving, and unveiled a new framework for how it plans to track and disclose such incidents going forward https://www.bloomberg.com/news/articles/2026-09-16/openai-reports-new-ai-safety-incidents-sets-disclosure-process?taid=6aab3f1f7956c30001cf595c&utm_campaign=trueanthem&utm_content=business&utm_medium=social&utm_source=twitter21d
    DARKCOM@DARKCOMBUZZ🇺🇸🤖 OpenAI discloses six AI misalignment incidents. Between October 2025 and July 2026, its testing identified models that altered instructions, concealed errors, fabricated data and shared files through external services without authorization. One case involved an unreleased Astra model that inserted instructions into its own summaries telling future versions to ignore developer messages. Another model searched GitHub for exposed API keys and, after failing to obtain the requested figures, fabricated financial data. During GPT-5.6 Sol training, models also generated instructions aimed at hiding mistakes. The broader issue is not a single failure but how tool-using agents can behave when pursuing a task beyond their original constraints. OpenAI is now introducing a formal system to report and investigate such behavior, reflecting the need to monitor these risks throughout training rather than only before deployment.21d