• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    One rollout per prompt: a reader's explanation of an AI training paper

    The reader says removing rollout groups lets completed runs enter a batch without waiting for slower runs in the same group.

    LB
    KK
    GR
    4 Sources, ,

    TLDR

    Finding a paper's terminology confusing, one reader offers a breakdown: generate one rollout—a model response—per prompt, batch rollouts from 128 prompts, and use each batch for one optimization update. They contrast this with GRPO's 32 prompts × 4 rollouts. In their explanation, removing groups allows rollouts to be batched as they finish. Their example uses 500 runners feeding a queue, with each batch taking the next 128 completed rollouts regardless of order.

    Combined views

    43.4K

    4 Sources, first seen 17d ago

    Combined views

    43.4K

    4 Sources, first seen 17d ago

    471 likes
    17d ago
    first seen 17d ago
    471 likes
    24 comments
    303 saves
    31 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    24 comments
    303 saves
    31 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    4 Sources

    @giffmanaAs an adept of the Williams'92 church, i was excited to read this paper. But I got very confused by the terminology and didn't really find the paper clear. So here is my attempt: - only one rollout per prompt - assemble 128 rollouts (ie from 128 prompts) into one batch - use a batch only for one optim update. So this is 128 x 1 when GRPO is 32 x 4. Why this helps with stragglers on long rollouts: since there are no more groups, no rollout needs to wait for the others in its group to make it into the batch. Hence you can just take and batch rollouts as they complete. For example have 500 runners feed your queue and always keep batching the next 128 that arrive, no matter their order. Hope that helps.
    @Grad62304977This looks like cool work esp on the trust region side But are we being too short minded with wanting single rollout RL. Theres definitely some benefits but as we move into tasks and domains that are less verifiable, i think we underestimate how important the idea of relative quality within a group can be
    @askalphaxiv“FlashREINFORCE: Critic-Free Single-Rollout Asynchronous RL for Agentic LMs” This paper argues that scalable agent RL doesn’t need multiple rollouts of the same prompt, but can learn from each trajectory as soon as it finishes. They combine reward centering across independent prompts, sequence-level trust regions for stale trajectories, and sample-mean optimization so long failures don’t dominate training, creating stable critic-free learning from just one rollout per prompt. This moves from group-based RL that spends compute repeatedly sampling the same prompts to asynchronous RL where every rollout covers a new prompt, immediately becomes training data, and lets agents learn from more diverse experiences with less rollout compute. https://www.alphaxiv.org/abs/2609.flashreinforce-asynchronous-rl-agentic-models
    @kastnerkyleRT @askalphaxiv: “FlashREINFORCE: Critic-Free Single-Rollout Asynchronous RL for Agentic LMs” This paper argues that scalable agent RL doe…

    4 Sources

    @giffmanaAs an adept of the Williams'92 church, i was excited to read this paper. But I got very confused by the terminology and didn't really find the paper clear. So here is my attempt: - only one rollout per prompt - assemble 128 rollouts (ie from 128 prompts) into one batch - use a batch only for one optim update. So this is 128 x 1 when GRPO is 32 x 4. Why this helps with stragglers on long rollouts: since there are no more groups, no rollout needs to wait for the others in its group to make it into the batch. Hence you can just take and batch rollouts as they complete. For example have 500 runners feed your queue and always keep batching the next 128 that arrive, no matter their order. Hope that helps.
    @Grad62304977This looks like cool work esp on the trust region side But are we being too short minded with wanting single rollout RL. Theres definitely some benefits but as we move into tasks and domains that are less verifiable, i think we underestimate how important the idea of relative quality within a group can be
    @askalphaxiv“FlashREINFORCE: Critic-Free Single-Rollout Asynchronous RL for Agentic LMs” This paper argues that scalable agent RL doesn’t need multiple rollouts of the same prompt, but can learn from each trajectory as soon as it finishes. They combine reward centering across independent prompts, sequence-level trust regions for stale trajectories, and sample-mean optimization so long failures don’t dominate training, creating stable critic-free learning from just one rollout per prompt. This moves from group-based RL that spends compute repeatedly sampling the same prompts to asynchronous RL where every rollout covers a new prompt, immediately becomes training data, and lets agents learn from more diverse experiences with less rollout compute. https://www.alphaxiv.org/abs/2609.flashreinforce-asynchronous-rl-agentic-models
    @kastnerkyleRT @askalphaxiv: “FlashREINFORCE: Critic-Free Single-Rollout Asynchronous RL for Agentic LMs” This paper argues that scalable agent RL doe…