Reactions from ranked influencers
17 postsVery big deal. Labs definitely have this in-house though Sol: «Directly extrapolating its local log-linear law to 1e25 produces impossible rewards above 100% and pushes RL toward approximately 60%, so that answer is unusable» but maybe not?
https://arxiv.org/abs/2607.16097 Scaling law for RL that connects rewards and pretraining loss, RL compute, on a synthetic (chess moves) setting. It also leads to the answer to the very interesting question, given fixed compute, what is the optimal allocation between pretraining and RL.
See also thread from @evangelinejy99
New Paper! As LLMs scale, where should compute go — stronger pretraining or more RL? Can we study pretraining and RL scaling jointly? We fit a pretraining–RL scaling law in a controlled chess testbed and use it to trace the optimal allocation. 🧵
https://arxiv.org/abs/2607.16097 Scaling law for RL that connects rewards and pretraining loss, RL compute, on a synthetic (chess moves) setting. It also leads to the answer to the very interesting question, given fixed compute, what is the optimal allocation between pretraining and RL.
We introduce a new testbed based on chess to study the pretraining-to-post-training scaling on a reasonable compute budget. It mimics the normal LLM training pipeline: PT on human games, SFT on synthetic reasoning traces, RL on chess puzzles.
https://arxiv.org/abs/2607.16097 Scaling law for RL that connects rewards and pretraining loss, RL compute, on a synthetic (chess moves) setting. It also leads to the answer to the very interesting question, given fixed compute, what is the optimal allocation between pretraining and RL.
very cool paper :)
New paper: Understanding Reasoning from Pretraining to Post-Training! We study the full LLM training pipeline from pretraining to post-training, find a joint scaling law, figure out how the compute should be allocated, and study what RL is doing to the policy. 🧵
Combined views
92.8K
17 posts, first seen 1d ago