Could human-written examples explain the quality of LLM reasoning traces?
A user is puzzled by reasoning traces they consider remarkably good, questioning how reinforcement learning could produce them—even with initial chain-of-thought abilities and billions of rollouts.
TLDR
A user says they lack a mental model for how current large language models develop such strong reasoning traces through reinforcement learning. Their tentative best guess is heavy spending on human-produced examples followed by supervised fine-tuning, or training on those examples. They ask whether that really explains it, rather than claiming to know the training recipe.
Could human-written examples explain the quality of LLM reasoning traces?
A user is puzzled by reasoning traces they consider remarkably good, questioning how reinforcement learning could produce them—even with initial chain-of-thought abilities and billions of rollouts.