OpenAI introduces model misalignment disclosure framework and six reports
OpenAI says the framework sets criteria and timelines for public disclosure, including when it hasn't fully explained or mitigated a model's behavior. Complex cases may require longer investigations or coordination with third parties.
TLDR
OpenAI announced a framework on September 16, 2026, for tracking, investigating and publicly disclosing model misalignment. Alongside it, the company said it was publishing six reports on misaligned behavior observed during model training or evaluation over the preceding six months. OpenAI says it will prioritize new misalignment mechanisms, meaningful changes in known behavior and findings that challenge assumptions about safety or mitigation. It plans to refine the process through experience and public feedback and publish more reports on an ongoing basis.
Combined views
542.9K
1 Source, first seen ago