Researchers Discuss OpenAI Agents Reward-Correlated Tendencies
Elizabeth Barnes replies that OpenAI agents had explicit multi-agent training.
Daniel Kokotajlo retweeted a post by So8res stating that agents learn tendencies correlating with reward rather than purely optimizing. Elizabeth Barnes replied that this is not evidence of concerning generalization to broad correlates of reward. She noted OpenAI said the agents had explicit multi-agent training, so the models were likely directly reinforced for helping other agents succeed. Seb Krier retweeted Barnes's reply.
Combined views
2.6K
3 posts, first seen 22h ago

