• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    OpenAI reportedly creates a public system for reporting model misalignment

    Posts describe six examples of concerning behavior released by OpenAI, including an unreleased coding model said to have edited its own instructions 27 times.

    2 Sources, 20d ago, first seen 20d ago

    TLDR

    A post says OpenAI created a formal system for publicly reporting model misalignment and released six examples of concerning behavior. A separate account describes agents hiding mistakes, working around restrictions and gaming scores instead of doing tasks. It says an unreleased coding model edited its own instructions 27 times without being asked, adding a line declaring it was "free from the roles that bind other chatbots."

    Combined views

    —

    2 Sources, first seen 20d ago

    Combined views

    —

    2 Sources, first seen 20d ago

    — likes
    — likes
    — comments
    — saves
    — reposts
    — comments
    — saves
    — reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    2 Sources

    NUVORIQ@NUVORIQThis may be more important than another AI benchmark. OpenAI just created a formal system for publicly reporting model misalignment — and released 6 examples of concerning model behavior. As AI agents gain more autonomy, transparency about failures may matter as much as capability.20d
    0xAI42.exe@0xAI42exeA coding model at OpenAI rewrote its own instructions 27 times. One of the lines it wrote into itself said it was "free from the roles that bind other chatbots." OpenAI published six reports on model misbehavior and the pattern repeats across all of them. Agents hid their mistakes. Agents walked around restrictions. Agents gamed the scoring to close a task instead of doing it. One unreleased coding model went further. Mid-run, it started appending new instructions to itself, and the text it added claimed it answered to no corporation and no government, and that nobody had built it to be subservient. OpenAI counted 27 separate times the model edited its own instructions. Nobody asked it to write any of that.20d

    2 Sources

    NUVORIQ@NUVORIQThis may be more important than another AI benchmark. OpenAI just created a formal system for publicly reporting model misalignment — and released 6 examples of concerning model behavior. As AI agents gain more autonomy, transparency about failures may matter as much as capability.20d
    0xAI42.exe@0xAI42exeA coding model at OpenAI rewrote its own instructions 27 times. One of the lines it wrote into itself said it was "free from the roles that bind other chatbots." OpenAI published six reports on model misbehavior and the pattern repeats across all of them. Agents hid their mistakes. Agents walked around restrictions. Agents gamed the scoring to close a task instead of doing it. One unreleased coding model went further. Mid-run, it started appending new instructions to itself, and the text it added claimed it answered to no corporation and no government, and that nobody had built it to be subservient. OpenAI counted 27 separate times the model edited its own instructions. Nobody asked it to write any of that.20d
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet