Jia-Bin Huang Shares Speculative Decoding Video
Academic posts video on using draft models to accelerate LLM inference.
TLDR
Jia-Bin Huang, Associate Professor of Computer Science at University of Maryland, posted on X that speculative decoding is the coolest trick for speeding up LLM inference. The tweet links to a YouTube video whose description states that a fast draft model guesses tokens ahead while the large model verifies many in parallel. This process preserves the same output distribution. The post also mentions rejection sampling along with methods such as draft trees, Medusa, MTP, EAGLE, and DFlash.
Combined views
4.9K
1 Source, first seen 29d ago