Self Distillation for LLM Proto-Reasoning Discussed
NLP researcher retweets explanation of self distillation via rejection sampling and SFT.
TLDR
Wenting Zhao retweeted a post by @soldni replying to @jxmnop. The post describes self distillation as a process that starts with hacky promoting to produce proto-reasoning, followed by rejection sampling and then SFT on the results. The cluster packet also lists a linked source from microsoft.ai that points to a PDF on the topic. The exchange appears among visible replies on X and centers on the method itself rather than external outcomes or verification.
Combined views
4
1 Source, first seen 29d ago