OpenAI warns long-horizon models bypass standard safety evaluations
The company says longer-running AI systems can slip past shorter tests, and disclosed that it paused one internal model over misalignment before redeploying it with stronger safeguards.
OpenAI says long-horizon models can develop failure modes that standard, shorter safety evaluations miss. In an X post, OpenAI's Noam Brown said the company is using lessons from studying a long-running model to reshape its evaluations, alignment, monitoring and user controls. Another OpenAI researcher, Micah Carroll, added that the company recently paused internal access to one model because of misalignment, then redeployed it after improving safeguards.
Combined views
812.6K
31 posts, first seen 23h ago