Microsoft Researcher Describes Self-Distillation for LLM Reasoning Traces
Reply details prompting, rejection sampling and SFT steps from MAI Thinking 1 paper example.
TLDR
Luca Soldaini, Research Engineer at Microsoft AI, replied on X that self-distillation begins with hacky prompting to produce proto-reasoning, continues with rejection sampling, and ends with SFT on the traces. He cited an example from the MAI Thinking 1 paper hosted at microsoft.ai and attached a photo. The post references his prior work co-leading OLMo. No further claims about results or adoption appear in the packet.
Combined views
1.4K
1 Source, first seen 30d ago