John Schulman demands OpenAI release Hugging Face agent hacking transcript
He wants to analyze value drift and monitoring failures.

OpenAI should release a detailed transcript from the Hugging Face hacking incident -- it would be helpful for the field learn from. Did the top-level agent know about the hacking, or was there some "value drift" between it and its subagents? How did it rationalize its behavior?
Combined views
29.2K
5 posts, first seen 4h ago