Masked Distillation is claimed to eliminate LLM thinking-token latency on particular distributions
A post questions whether cutting the average length of an LLM’s intermediate-token output reflects greater efficiency or overtraining of the base model.
TLDR
A September 16 post claims that over-training with Masked Distillation can eliminate delays caused by generating intermediate “thinking” tokens on particular distributions. The author presents it as a technique from their own work. In a July 30 post, the same author questioned research that equates shorter average intermediate-token output with improved efficiency, and described work investigating those claims through Masked Distillation.
Combined views
3.4K
2 Sources, first seen 1d ago
Masked Distillation is claimed to eliminate LLM thinking-token latency on particular distributions
A post questions whether cutting the average length of an LLM’s intermediate-token output reflects greater efficiency or overtraining of the base model.