Report
New on-policy distillation method is claimed to help student models match or exceed teachers
HuggingPapers says the method extrapolates RL-induced representation residuals, with code and checkpoints planned for release.
TLDR
HuggingPapers describes a new on-policy distillation method that extrapolates RL-induced representation residuals. It says student models can match or exceed their teachers across four model pairs, and that code and checkpoints are planned for release.
Combined views
8.1K
2 Sources, first seen 20h ago
likes
