OpenAI releases framework for reporting model misalignment
A user sharing the release says it includes six reports of “unexpected or concerning” model behavior and defines processes for tracking, investigating and disclosing misalignment.
TLDR
OpenAI has published a framework for reporting model misalignment, according to a user sharing the release. It covers tracking, investigation and disclosure, alongside six reports of unexpected or concerning model behavior. The user argues that a common reporting format could make comparisons across labs and outside safety analysis easier.
Combined views
—
1 Source, first seen ago