Unsloth is one tester’s pick among Qwen3.8-27B NVFP4 variants
The tester reports very close accuracy across variants, but says NVIDIA’s MTP-4 was roughly 2.5–3× slower than Unsloth in long-context coding tests on an RTX Pro 6000.
TLDR
One tester says Qwen3.8-27B NVFP4 variants are very close in accuracy, with memory use and speed setting them apart. They call minima-ai’s version a good choice for limited GPU memory, but note that it lacks multi-token prediction (MTP) for faster inference.
In their long-context coding tests on an RTX Pro 6000, NVIDIA’s version with MTP-4 was about 2× faster than without MTP, yet roughly 2.5–3× slower than Unsloth. Their pick was Unsloth NVFP4.
Combined views
108
1 Source, first seen 17d ago