OpenAI releases model misalignment disclosure framework and six reports
OpenAI says the framework sets criteria and timelines for public disclosure, including when it has not yet fully explained or mitigated the behavior.
TLDR
OpenAI announced the framework on September 16, 2026, alongside six reports on misaligned behavior it says it observed during model training or evaluation in the preceding six months. The company says it will prioritize new misalignment mechanisms, meaningful changes in known behavior and findings that challenge assumptions about safety or mitigation. It notes that complex cases may require longer investigation or coordination with third parties. OpenAI plans to refine the process through experience and public feedback and publish more reports.
Combined views
61.7K
4 Sources, first seen ago