SemiAnalysis Posts on vLLM Kimi Throughput Gains
SemiAnalysis account notes vLLM gains on Kimi model and introduces AgentX benchmark.
TLDR
The official SemiAnalysis account posted replies crediting vLLM optimizations with raising high-concurrency throughput by more than 6x in 2 weeks for the Kimi K3 2.8T model. The posts reference a chart of token throughput per chip versus P90 interactivity and state that the firm's AgentX benchmark uncovered issues in multi-turn long-context tasks. The linked article is titled AgentX - InferenceXv3: Does CUDA Moat Hold up in Agentic Inferencing? It lists a $3 Million USD dataset open sourced along with 1 Mil+ Context Length, Multiturn, Sub Agents, 95%+ KVCache HitRate, GB300 NVL72, MI355, and B200.
Combined views
14.4K
2 Sources, first seen 27d ago