OpenAI Pauses Frontier RL Training Over Misalignment
Company cites misalignment signals and growing risks in advanced models.
TLDR
OpenAI announced it temporarily paused reinforcement learning training on its latest models intended for deployment. The company said it used the time to harden research environments, red-team them, and expand monitoring coverage. Sam Altman told Alex Heath that unreleased models showed various degrees of misalignment. Greg Brockman noted the slowdown included the largest planned frontier RL run. Altman added that the pause affects further-out releases while the company still expects to ship new models soon.