Inco AI Releases DFlash 2 for Faster Qwen Inference
Researchers react to the speculative decoding update for Qwen models on Apple hardware.
TLDR
Zhijian Liu posted that Inco AI released DFlash 2, the next version of a speculative decoding system first developed at Z Lab. The post states Qwen3.8-27B reaches 70 tokens per second on an M5 Max MacBook Pro, up to 4.6 times faster than standard autoregressive decoding while producing the same output. Inco AI called it the company's first release. Academics Beidi Chen, Lianhui Qin, Song Han, and others replied with praise for the speed gain on the MacBook Pro hardware.
Combined views
1.1M
10 Sources, first seen 43d ago
