Tester favors Unsloth for speed among Qwen3.8-27B NVFP4 variants
In one user's long-context coding tests on an RTX Pro 6000, NVIDIA's version with MTP-4 was roughly 2.5–3× slower than Unsloth, despite being about 2× faster than running without MTP.
TLDR
A user comparing Qwen3.8-27B NVFP4 variants says they are very close in accuracy, with memory use and speed setting them apart. They recommend minima-ai/mnma_qwen3.8_27b_nvfp4 when GPU memory is limited, but note that it lacks multi-token prediction (MTP) for faster inference. Their overall pick is Unsloth: in long-context coding tests on an RTX Pro 6000, they report NVIDIA's MTP-4 was about 2× faster than no MTP but still roughly 2.5–3× slower than Unsloth.
Combined views
7.2K
1 Source, first seen 17d ago