Users are excited about LLMs improving reasoning by reading consensus instead of rewards, seeing it as a promising new method.
Based on 1 visible X reactions from 4 accounts; directional sample.
Ask a question below.
Published answers will appear here.
@j0hngou love seeing this.
Consensus is usually compressed into a vote or scalar reward. This work instead uses a consensus solution as context for a frozen self-teacher, turning agreement into dense token-level supervision. Recovers much of the benefit of gold solutions, without labels
@iatitov Hard "E-M"?
Users are excited about LLMs improving reasoning by reading consensus instead of rewards, seeing it as a promising new method.
Based on 1 visible X reactions from 4 accounts; directional sample.
Ask a question below.
Published answers will appear here.