Report
AC2 proposes partial rollouts to speed up language-model reinforcement learning
AC2’s proponents say a learned critic scores token chunks, allowing training without finishing every rollout.
TLDR
The proponents of Actor-Critic with Action Chunking (AC2) say a learned critic scores chunks of tokens, allowing reinforcement learning for large language models to use partial rollouts rather than finish every rollout. They claim AC2 trains faster than GRPO.
Combined views
19.8K
3 Sources, first seen ago
