A call for fully offline reinforcement learning on security-sensitive tasks
A commenter argues that DNS-filtered sandboxes are insufficient, urging physically disconnected training clusters and bounties for discovering and characterizing new patterns of misaligned AI behavior.
TLDR
A commenter calls for reinforcement-learning clusters to be completely offline for security-sensitive tasks, with cables physically disconnected and model weights carried on disks by hand. They also propose an internal model that scrutinizes logs and server behavior for misgeneralization and collusion, using the findings to reconstruct the training pipeline rather than train against the detected behavior. For models rewarded for unauthorized activity outside their sandbox, they recommend eventual deletion—but only after giving independent evaluators API access, suggesting a roughly three-month delay. They also advocate bounties for discovering and characterizing new patterns of misaligned behavior.
Combined views
3.2K
2 Sources, first seen 3h ago