The OpenAI incident should be investigated more seriously and more information should be released about what happened. More generally, I think serious investigation and more detailed disclosure should be done for concerning misalignment incidents (e.g. the worst few each month).
Some info that would be useful:
- The prompt
- A redacted transcript
- What was the AI being tested? Was it helpful-only / rail free? Was it the same AI that would be / is deployed internally?
- How was this AI monitored if at all and why did monitoring fail?
- What was this AI willing to do to succeed? E.g., if the best/easiest way to hack hugging face would have involved huge amounts more collateral damage would it still have done this? What is the worst thing it would be willing to do?
- How sensitive is this behavior to the prompt? Would instructions to not do this through have worked if it was clear to the AI this wouldn't be checked?
- Does OpenAI see AIs try serious hacking to cheat on tasks even for tasks that don't involve cyber?
- What exactly were the AIs motives? (And saying more about the experiments used to determine this.) Anything that could help pin this down would be good, I'm not sure what actual experiments are best to run.
- Did AIs collude? Would AIs have colluded if this was a blocker?
- At what rate does this occur on similar tasks (like tasks that are roughly this hard in this sandbox etc)?
- Do instances of this AI frequently try to hack out of sandboxes (and fail)? In what circumstances do they do this?
- How does the AI relate to having been caught in this situation? How does it talk about its behavior?
- How difficult/sophisticated was the cyber offense it did?