Users express optimism about OpenAI's safety alignment details for long-horizon models because the issues seem relatively fixable.
Based on 1 visible X reactions from 1 accounts; directional sample.
Ask a question below.
Published answers will appear here.
A few more thoughts on why / how this seems relatively fixable to me:
Just read this post, which I appreciate for sharing more details -- seems like long-horizon instruction following of user constraints was a weakness. But it unfortunately still doesn't say very much about RL training incentives. https://openai.com/index/safety-alignment-long-horizon-models/
Users express optimism about OpenAI's safety alignment details for long-horizon models because the issues seem relatively fixable.
Based on 1 visible X reactions from 1 accounts; directional sample.
Ask a question below.
Published answers will appear here.