SGLang adds 4-bit NVFP4 KV cache support for Nvidia Blackwell
LMSYS says the cache uses about 56% of FP8's memory footprint per token and boosts long-context decoding throughput by up to 78%.
TLDR
LMSYS announced NVFP4 KV cache support in SGLang for Nvidia Blackwell, developed with Alibaba's Qwen team and Nvidia. It reports decoding throughput gains of 37%, 58% and 78% at context lengths of 32K, 160K and 1M, respectively. LMSYS also says accuracy matched FP8 on GPQA-Diamond and AIME 2025 using Qwen3.5-397B-A17B. The cache can be enabled in SGLang with the flag --kv-cache-dtype nvfp4.
