Announcement
SGLang brings Kimi K3 inference optimizations to NVIDIA Vera Rubin
The SGLang team says MoE tail fusion delivered a 5.9% end-to-end inference speedup on early-access Rubin hardware.
TLDR
SGLang says it worked with NVIDIA on early-access Vera Rubin hardware to accelerate Kimi K3 inference. It reports up to 20% faster FP8 MLA at batch 1 and 128K context, 20% faster KDA verification with bitwise-identical output, and a 5.9% end-to-end inference speedup from MoE tail fusion. SGLang also says it powers Miles' reinforcement-learning rollouts on Rubin, including agentic RL with 64 concurrent sandboxes on the Vera CPU.
Combined views
3.4K
3 Sources, first seen ago