Masked Distillation is claimed to eliminate LLM thinking-token latency on particular distributions · Digg