Announcement
Slime v0.4.0 adds distributed rollouts and independent recovery for RL training
The slime team says the update pairs synchronous training for algorithmic exploration with fully asynchronous training for throughput at scale.
TLDR
The slime team says OpenAI’s Navier-Stokes effort used about 10,000 concurrent agents and 130 billion output tokens, prompting it to rethink the scale needed for RL infrastructure. Slime v0.4.0 spreads rollout orchestration across cluster CPUs, adds straw middleware for durable queues and shared tensor storage, and lets Megatron restart while healthy SGLang engines keep running, preserving accepted rollout work.
Combined views
5K
1 Source, first seen ago