• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    SemiAnalysis Posts on vLLM Kimi Throughput Gains

    SemiAnalysis account notes vLLM gains on Kimi model and introduces AgentX benchmark.

    SE
    2 Sources, 27d ago, first seen 27d ago

    TLDR

    The official SemiAnalysis account posted replies crediting vLLM optimizations with raising high-concurrency throughput by more than 6x in 2 weeks for the Kimi K3 2.8T model. The posts reference a chart of token throughput per chip versus P90 interactivity and state that the firm's AgentX benchmark uncovered issues in multi-turn long-context tasks. The linked article is titled AgentX - InferenceXv3: Does CUDA Moat Hold up in Agentic Inferencing? It lists a $3 Million USD dataset open sourced along with 1 Mil+ Context Length, Multiturn, Sub Agents, 95%+ KVCache HitRate, GB300 NVL72, MI355, and B200.

    Combined views

    14.4K

    2 Sources, first seen 27d ago

    Combined views

    14.4K

    2 Sources, first seen 27d ago

    37 likes
    37 likes
    3 comments
    8 saves
    5 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    3 comments
    8 saves
    5 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    @SemiAnalysis_@vllm_project With these and other optimizations, the team at vLLM brought Kimi inference to new heights, increasing high-concurrency throughput by more than 6x in 2 weeks. (4/5)

    2 Sources

    @SemiAnalysis_@vllm_project With these and other optimizations, the team at vLLM brought Kimi inference to new heights, increasing high-concurrency throughput by more than 6x in 2 weeks. (4/5)