Reaction
The case for calibrating smaller AI models after distillation
A user argues that smaller models distilled from larger ones need follow-up reinforcement learning to judge their own capabilities.
TLDR
A user argues that distilling a general-purpose smaller model from a larger model trained with reinforcement learning is not enough. Without its own follow-up reinforcement learning for calibration, they say, the smaller model will be overconfident about its capabilities and use too little reasoning.
Combined views
3.4K
1 Source, first seen ago