Mercor Releases Open RL Guide for 397B Qwen Model
Guide for agentic RL post-training of 397B Qwen model uses SkyRL and 2000 tasks.
Robert Nishihara posted about an agentic RL training guide for the Qwen 3.5 397B model created with SkyRL. The materials, linked from a tweet by edwardjhu, are presented as coming from Mercor Research and include training scripts for applying DPPO. The post states the approach yields very large improvements when trained on roughly 2000 expert-labeled tasks. No further confirmation or independent details appear in the packet.
Combined views
74.2K
5 posts, first seen 23h ago