Kimi K3 is reportedly 2.2–2.8× faster on vLLM
Inferact says it co-led the optimization effort with Red Hat AI, Nvidia and Huawei.
TLDR
Inferact reported a 2.2–2.8× speedup for Kimi K3 on vLLM in its September 18 announcement. It says the work spans scheduling, KDA state handling and custom mixture-of-experts (MoE) kernels, and shared a technical deep dive into the optimizations.
Combined views
1.8K
2 Sources, first seen 4h ago
Kimi K3 is reportedly 2.2–2.8× faster on vLLM
Inferact says it co-led the optimization effort with Red Hat AI, Nvidia and Huawei.