DeepSeek V4 Flash’s decisions reportedly distilled into a 4B model
A user says the smaller model beat its teacher’s instant-answer mode at one-twentieth the size, taking about 22 milliseconds per decision.
TLDR
A user reports spending 26 hours on a DGX Spark distilling DeepSeek V4 Flash’s decisions into a 4-billion-parameter model. They describe the teacher model as having 157GB of weights and say the smaller model surpassed its teacher’s instant-answer mode at one-twentieth the size, with each decision taking about 22 milliseconds.
Combined views
414.4K
2 Sources, first seen 1d ago
DeepSeek V4 Flash’s decisions reportedly distilled into a 4B model
A user says the smaller model beat its teacher’s instant-answer mode at one-twentieth the size, taking about 22 milliseconds per decision.