• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    UnslothAI Announces MTP Speedup for Qwen3.8-Flash

    Retweet by Unsloth AI co-founder shares local speed claims for Qwen model.

    DH
    UA
    2 Sources, 28d ago, first seen 28d ago

    TLDR

    Unsloth AI posted that Qwen3.8-Flash can now run 1.7× faster locally with MTP. The post states GGUFs can reach 170 tokens/s on an RTX PRO 6000. The company states MTP enables Qwen3.8-Flash-Next inference 1.3–1.7× faster with no accuracy change. Daniel Han retweeted the announcement. Links point to GGUF files on Hugging Face and a local-run guide on unsloth.ai. The two posts in the visible conversation supply the claims. No independent confirmation appears in the packet.

    Combined views

    97.1K

    2 Sources, first seen 28d ago

    Combined views

    97.1K

    2 Sources, first seen 28d ago

    1.1K likes
    1.1K likes
    59 comments
    513 saves
    136 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    59 comments
    513 saves
    136 reposts
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    2 Sources

    @UnslothAIQwen3.8-Flash can now run 1.7× faster locally with MTP!⚡️ GGUFs can reach 170 tokens/s on a RTX PRO 6000. MTP enables Qwen3.8-Flash-Next ~1.3–1.7× faster inference with no accuracy change. GGUFs: https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF Guide: https://unsloth.ai/docs/models/qwen3.8-next
    @danielhanchenRT @UnslothAI: Qwen3.8-Flash can now run 1.7× faster locally with MTP!⚡️ GGUFs can reach 170 tokens/s on a RTX PRO 6000. MTP enables Qwen…

    2 Sources

    @UnslothAIQwen3.8-Flash can now run 1.7× faster locally with MTP!⚡️ GGUFs can reach 170 tokens/s on a RTX PRO 6000. MTP enables Qwen3.8-Flash-Next ~1.3–1.7× faster inference with no accuracy change. GGUFs: https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF Guide: https://unsloth.ai/docs/models/qwen3.8-next
    @danielhanchenRT @UnslothAI: Qwen3.8-Flash can now run 1.7× faster locally with MTP!⚡️ GGUFs can reach 170 tokens/s on a RTX PRO 6000. MTP enables Qwen…