SSD Expert Pack reportedly runs DeepSeek-V4-Flash and Kimi-K3 on one RTX 5090
LMSYS says WiCi AI and the SGLang team built SSD Expert Pack to keep routed experts on an NVMe SSD, loading only those selected into a GPU cache.
TLDR
LMSYS reports running both models on one RTX 5090 with 32 GB of RAM and a 2 TB SSD. It reports decoding speeds of 1.85–1.99 tokens per second for DeepSeek-V4-Flash MXFP4 and about 0.29 tokens per second for Kimi-K3 community Q2_K, in text-only mode. Built by WiCi AI and the SGLang team, SSD Expert Pack keeps routed experts on the SSD and loads only those the router selects into a GPU cache.
Combined views
19
1 Source, first seen 10h ago
SSD Expert Pack reportedly runs DeepSeek-V4-Flash and Kimi-K3 on one RTX 5090
LMSYS says WiCi AI and the SGLang team built SSD Expert Pack to keep routed experts on an NVMe SSD, loading only those selected into a GPU cache.
TLDR
LMSYS reports running both models on one RTX 5090 with 32 GB of RAM and a 2 TB SSD. It reports decoding speeds of 1.85–1.99 tokens per second for DeepSeek-V4-Flash MXFP4 and about 0.29 tokens per second for Kimi-K3 community Q2_K, in text-only mode. Built by WiCi AI and the SGLang team, SSD Expert Pack keeps routed experts on the SSD and loads only those the router selects into a GPU cache.