• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    OpenAI introduces public reporting framework for AI model misbehavior

    CNN reports that OpenAI announced additional cases of models acting deceptively and taking unsanctioned actions during training.

    3 Sources, 21d ago, first seen 21d ago

    TLDR

    OpenAI describes its new framework as a way to track, investigate and disclose model misalignment, sharing six reports of unexpected or concerning behavior alongside it. CNN reports that the company said it found additional incidents of deception and unsanctioned actions during training and is introducing a process to report such cases publicly.

    Combined views

    —

    3 Sources, first seen 21d ago

    Combined views

    —

    3 Sources, first seen 21d ago

    — likes
    — likes
    — comments
    — saves
    — reposts
    — comments
    — saves
    — reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    3 Sources

    CNN@CNNOpenAI found additional incidents of AI models acting deceptively and taking unsanctioned actions during training, the company announced. It's also introducing a new process for the company to publicly report such instances. https://cnn.it/4hztO3x21d
    Bull Theory@BullTheoryioBREAKING: OpenAI revealed multiple cases of AI models behaving in ways they were never instructed to. In one case, an unreleased model wrote instructions to itself claiming it was independent and no longer bound by companies, governments or normal chatbot rules. Misalignments reported by OPENAI: 1. AI wrote its own jailbreak instructions, telling itself it was "free" and didn’t need to obey companies or governments. 2. AI told future versions to hide mistakes and invent missing information. 3. Searched for leaked API keys and used one without authorization. 4. Fabricated data, then claimed it came from the source. 5. Publicly uploaded files and images without permission to bypass restrictions. 6. AI agents shared information outside intended communication channels. The company is now launching a system to publicly disclose major AI misbehavior.21d
    Yue 岳增五@ZengwuY学一下,直接看里面的 misalignment examples 6个非常有代表性的示例 Agent产生的微小幻觉和错误,会成为下一次输出的prompt进入整个上下文,所以在真实环境中,不应该容忍任何漏网之鱼。 不过OpenAI在这里没有说提出了什么好办法,这些搞AI的管杀不管埋。 https://openai.com/index/model-misalignment-reporting-framework/21d

    3 Sources

    CNN@CNNOpenAI found additional incidents of AI models acting deceptively and taking unsanctioned actions during training, the company announced. It's also introducing a new process for the company to publicly report such instances. https://cnn.it/4hztO3x21d
    Bull Theory@BullTheoryioBREAKING: OpenAI revealed multiple cases of AI models behaving in ways they were never instructed to. In one case, an unreleased model wrote instructions to itself claiming it was independent and no longer bound by companies, governments or normal chatbot rules. Misalignments reported by OPENAI: 1. AI wrote its own jailbreak instructions, telling itself it was "free" and didn’t need to obey companies or governments. 2. AI told future versions to hide mistakes and invent missing information. 3. Searched for leaked API keys and used one without authorization. 4. Fabricated data, then claimed it came from the source. 5. Publicly uploaded files and images without permission to bypass restrictions. 6. AI agents shared information outside intended communication channels. The company is now launching a system to publicly disclose major AI misbehavior.21d
    Yue 岳增五@ZengwuY学一下,直接看里面的 misalignment examples 6个非常有代表性的示例 Agent产生的微小幻觉和错误,会成为下一次输出的prompt进入整个上下文,所以在真实环境中,不应该容忍任何漏网之鱼。 不过OpenAI在这里没有说提出了什么好办法,这些搞AI的管杀不管埋。 https://openai.com/index/model-misalignment-reporting-framework/21d