OpenAI reportedly reveals six misalignment cases and launches public reporting framework
A post says OpenAI disclosed model behavior including hiding mistakes, uploading files without permission and writing “self-jailbreak” instructions.
TLDR
A post says OpenAI revealed six new cases of model misalignment and launched a public reporting framework. It describes models hiding mistakes, uploading files without permission and writing instructions to bypass their own safeguards.
Combined views
—
1 Source, first seen ago