OpenAI releases model misalignment disclosure framework and six reports
OpenAI says the framework sets disclosure criteria and timelines even for behavior it hasn't fully explained or mitigated. Complex cases may require longer investigations or coordination with third parties.
TLDR
OpenAI announced the framework on September 16, 2026, alongside six reports on misaligned behavior it observed during model training or evaluation over the previous six months. The company says it will prioritize cases that reveal new misalignment mechanisms, meaningful changes in known behavior, or findings that challenge assumptions about safety or mitigation. It plans to publish more reports on an ongoing basis and refine the process through experience and public feedback.
Combined views
542.9K
1 Source, first seen 14d ago