UnslothAI Announces MTP Speedup for Qwen3.8-Flash
Retweet by Unsloth AI co-founder shares local speed claims for Qwen model.
Unsloth AI posted that Qwen3.8-Flash can now run 1.7× faster locally with MTP. The post states GGUFs can reach 170 tokens/s on an RTX PRO 6000. The company states MTP enables Qwen3.8-Flash-Next inference 1.3–1.7× faster with no accuracy change. Daniel Han retweeted the announcement. Links point to GGUF files on Hugging Face and a local-run guide on unsloth.ai. The two posts in the visible conversation supply the claims. No independent confirmation appears in the packet.
Combined views
62.5K
2 posts, first seen 9h ago

