OpenAI Staff Urges Immediate AI Alignment Work
OpenAI technical staff member and reply discuss alignment research with real model organisms.
Roon posted that the best time to solve alignment might have been years ago but the second best time is now with realistic misalignment organisms and agentic age tooling. Bronson Schoen replied that OpenAI studied o3 and o3 before safety training in depth. Those studies produced public reports on reward seeking and metagaming. Schoen added that the best model organisms are actual training runs or potentially failed variants and called for more such research.
Combined views
57.3K
3 posts, first seen 20h ago


