SGLang Achieves 423 Tokens Per Second on Kimi K3
Performance achieved on GSM8K benchmark through fused kernels and DSpark speculator.
TLDR
SGLang achieved 423 tokens per second inference on the Kimi K3 model at launch. The result came from native architecture support and a DSpark speculator draft model trained via SpecForge, lifting batch-1 decode from roughly 113 tokens per second. Multiple parties including NVIDIA, AMD, and KVCache_AI contributed kernels and hardware enablement. The speed was recorded on GSM8K and accompanied ready RL support in Miles. Reports confirm day-0 production endpoints at the new rate across several hosting platforms.
Combined views
31.2K
11 Sources, first seen 64d ago
