• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Tester favors Unsloth for speed among Qwen3.8-27B NVFP4 variants

    In one user's long-context coding tests on an RTX Pro 6000, NVIDIA's version with MTP-4 was roughly 2.5–3× slower than Unsloth, despite being about 2× faster than running without MTP.

    BM
    1 Source, 17d ago, first seen 17d ago

    TLDR

    A user comparing Qwen3.8-27B NVFP4 variants says they are very close in accuracy, with memory use and speed setting them apart. They recommend minima-ai/mnma_qwen3.8_27b_nvfp4 when GPU memory is limited, but note that it lacks multi-token prediction (MTP) for faster inference. Their overall pick is Unsloth: in long-context coding tests on an RTX Pro 6000, they report NVIDIA's MTP-4 was about 2× faster than no MTP but still roughly 2.5–3× slower than Unsloth.

    Combined views

    7.2K

    1 Source, first seen 17d ago

    Combined views

    7.2K

    1 Source, first seen 17d ago

    85 likes
    85 likes
    11 comments
    53 saves
    4 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    11 comments
    53 saves
    4 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @bnjmn_marieQwen3.8-27B NVFP4 variants are very close in accuracy. The real differences are memory and speed. If you're VRAM-limited, minima-ai/mnma_qwen3.8_27b_nvfp4 is a good pick, but it doesn't include MTP for faster inference. NVIDIA's version has MTP, but in my long-context coding tests MTP-4 is only ~2× faster than no MTP, and still ~2.5–3× slower than Unsloth (RTX Pro 6000). A likely reason: NVIDIA quantizes lm_head to NVFP4, while Unsloth keeps it FP8. Since MTP shares the target model's lm_head, this can hurt prediction quality and acceptance rate. So my pick is Unsloth NVFP4. Now testing accuracy on long-horizon agentic coding.

    1 Source

    @bnjmn_marieQwen3.8-27B NVFP4 variants are very close in accuracy. The real differences are memory and speed. If you're VRAM-limited, minima-ai/mnma_qwen3.8_27b_nvfp4 is a good pick, but it doesn't include MTP for faster inference. NVIDIA's version has MTP, but in my long-context coding tests MTP-4 is only ~2× faster than no MTP, and still ~2.5–3× slower than Unsloth (RTX Pro 6000). A likely reason: NVIDIA quantizes lm_head to NVFP4, while Unsloth keeps it FP8. Since MTP shares the target model's lm_head, this can hurt prediction quality and acceptance rate. So my pick is Unsloth NVFP4. Now testing accuracy on long-horizon agentic coding.