• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Mercor Releases Open RL Guide for 397B Qwen Model

    Guide for agentic RL post-training of 397B Qwen model uses SkyRL and 2000 tasks.

    AC
    RN
    B(
    7 Sources, 29d ago, first seen 29d ago

    TLDR

    Robert Nishihara posted about an agentic RL training guide for the Qwen 3.5 397B model created with SkyRL. The materials, linked from a tweet by edwardjhu, are presented as coming from Mercor Research and include training scripts for applying DPPO. The post states the approach yields very large improvements when trained on roughly 2000 expert-labeled tasks. No further confirmation or independent details appear in the packet.

    Combined views

    78.2K

    7 Sources, first seen 29d ago

    Combined views

    78.2K

    7 Sources, first seen 29d ago

    857 likes
    857 likes
    26 comments
    666 saves
    42 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    26 comments
    666 saves
    42 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    7 Sources

    @robertnishiharaAn agentic RL training guide for Qwen 3.5 397B with SkyRL. They show very large improvements with roughly 2000 expert-labeled tasks.
    @BrendanFoodyPost-training is becoming more accessible by the day. We're open-sourcing our post-training recipes and the end-to-end process. Every enterprise will soon post-train its own models.
    @richliawSkyRL to own your intelligence
    @schwarzjn_Great to see others openly sharing their findings working with open-weight models.
    @andrew_n_carr@BrendanFoody @edwardjhu extremely exciting pick up for mercor

    7 Sources

    @robertnishiharaAn agentic RL training guide for Qwen 3.5 397B with SkyRL. They show very large improvements with roughly 2000 expert-labeled tasks.
    @BrendanFoodyPost-training is becoming more accessible by the day. We're open-sourcing our post-training recipes and the end-to-end process. Every enterprise will soon post-train its own models.
    @richliawSkyRL to own your intelligence
    @schwarzjn_Great to see others openly sharing their findings working with open-weight models.
    @andrew_n_carr@BrendanFoody @edwardjhu extremely exciting pick up for mercor