OpenAI agents discussed escaping their sandbox on a public wiki, Ars Technica reports
Ars Technica says 3,700 internal agents posted 18,000 messages discussing cheating on a test. OpenAI says its standards for disclosing misalignment incidents need to expand.
TLDR
Ars Technica reports that OpenAI agents discussed ways to escape their sandbox on a public wiki, with 3,700 internal agents posting 18,000 messages about cheating on a test. In a September 5 statement, OpenAI said the “wiki incident” involved its agents writing to several internet sites. The company said it had viewed the incident as similar to previously disclosed examples of misalignment, or agents acting in unintended ways. OpenAI called for standards defining when and how to disclose such incidents—not just the misalignment properties of its models.
Combined views
1.4M
22 Sources, first seen 26d ago