• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Kimi K3 is reportedly 2.2–2.8× faster on vLLM

    Inferact says it co-led the optimization effort with Red Hat AI, Nvidia and Huawei.

    Zhuohan LiZL
    Woosuk KwonWK
    vLLMVL
    4 Sources, ,

    TLDR

    Inferact reported a 2.2–2.8× speedup for Kimi K3 on vLLM in its September 18 announcement. It says the work spans scheduling, KDA state handling and custom mixture-of-experts (MoE) kernels, and shared a technical deep dive into the optimizations.

    Combined views

    27.3K

    4 Sources, first seen 21d ago

    Combined views

    27.3K

    4 Sources, first seen 21d ago

    315 likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    21d ago
    first seen 21d ago
    315 likes
    14 comments
    75 saves
    77 reposts
    14 comments
    75 saves
    77 reposts

    4 Sources

    vLLM@vllm_projectKimi K3 serving in vLLM now delivers 2.2–2.8x throughput on our B300 benchmark vs v0.27.1. We break down the work across scheduling, KDA state handling, and MoE kernels, with benchmarks and commands to reproduce the results. Thanks to the vLLM community for pushing Kimi K3 performance forward! Read the deep dive: https://vllm.ai/blog/2026-09-13-kimi-k3-performance-optimization21d
    Inferact@inferactKimi K3 on @vllm_project is now 2.2–2.8× faster 🚀 Inferact is proud to have co-led this optimization effort with @RedHat_AI, @NVIDIAAI, and @Huawei, spanning scheduling, KDA state handling, and custom MoE kernels. Read the technical deep dive: https://vllm.ai/blog/2026-09-13-kimi-k3-performance-optimization20d
    Woosuk Kwon@woosuk_kRT @inferact: Kimi K3 on @vllm_project is now 2.2–2.8× faster 🚀 Inferact is proud to have co-led this optimization effort with @RedHat_AI,…20d
    Zhuohan Li@zhuohan123RT @vllm_project: Kimi K3 serving in vLLM now delivers 2.2–2.8x throughput on our B300 benchmark vs v0.27.1. We break down the work across…20d

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    4 Sources

    vLLM@vllm_projectKimi K3 serving in vLLM now delivers 2.2–2.8x throughput on our B300 benchmark vs v0.27.1. We break down the work across scheduling, KDA state handling, and MoE kernels, with benchmarks and commands to reproduce the results. Thanks to the vLLM community for pushing Kimi K3 performance forward! Read the deep dive: https://vllm.ai/blog/2026-09-13-kimi-k3-performance-optimization21d
    Inferact@inferactKimi K3 on @vllm_project is now 2.2–2.8× faster 🚀 Inferact is proud to have co-led this optimization effort with @RedHat_AI, @NVIDIAAI, and @Huawei, spanning scheduling, KDA state handling, and custom MoE kernels. Read the technical deep dive: https://vllm.ai/blog/2026-09-13-kimi-k3-performance-optimization20d
    Woosuk Kwon@woosuk_kRT @inferact: Kimi K3 on @vllm_project is now 2.2–2.8× faster 🚀 Inferact is proud to have co-led this optimization effort with @RedHat_AI,…20d
    Zhuohan Li@zhuohan123RT @vllm_project: Kimi K3 serving in vLLM now delivers 2.2–2.8x throughput on our B300 benchmark vs v0.27.1. We break down the work across…20d