• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Announcement

    Slime v0.4.0 adds distributed rollouts and independent recovery for RL training

    The slime team says the update pairs synchronous training for algorithmic exploration with fully asynchronous training for throughput at scale.

    slimeSL
    1 Source, 2h ago, first seen 2h ago

    TLDR

    The slime team says OpenAI’s Navier-Stokes effort used about 10,000 concurrent agents and 130 billion output tokens, prompting it to rethink the scale needed for RL infrastructure. Slime v0.4.0 spreads rollout orchestration across cluster CPUs, adds straw middleware for durable queues and shared tensor storage, and lets Megatron restart while healthy SGLang engines keep running, preserving accepted rollout work.

    Combined views

    5K

    1 Source, first seen 2h ago

    Combined views

    5K

    1 Source, first seen 2h ago

    68 likes
    68 likes
    4 comments
    39 saves
    7 reposts
    4 comments
    39 saves
    7 reposts
    Featured Source

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    slime@slime_frameworkOpenAI's Navier-Stokes effort used ~10,000 concurrent agents and ~130B output tokens. It made us rethink the scale ahead for RL infrastructure. slime v0.4.0 is a first step toward greater scale, with four architectural changes: • Focus: synchronous training for algorithmic exploration and correctness; fully async training for throughput at scale. • Distributed rollout: spread orchestration across cluster CPUs, with independent generation pools and event loops. • Persistent storage: introduce straw, our new storage middleware for rollout and training data, providing durable queues and shared tensor storage. • Independent recovery: restart Megatron while healthy SGLang engines stay running, preserving accepted rollout work. More on the architecture below 👇 https://github.com/THUDM/slime/releases/tag/v0.4.0 1/52h