SGLang v0.5.20 adds Intel XPU to standard releases
The SGLang project says the release brings up to 52% faster decoding with reinforcement-learning sampling masks and up to 12.5× faster ROCm model loading.
TLDR
SGLang announced v0.5.20 with Intel XPU joining its standard releases. The project says reinforcement-learning sampling masks improve rollout reliability and decoding speed, while SGLang-Diffusion gets up to about 38% lower end-to-end latency. Other highlights include SGLang Simulator for scheduler and cache experiments on CPUs, plus new models including GLM-5.3-Flash, Qwen3.8-Flash-Next and K2 Horizon.
SGLang v0.5.20 adds Intel XPU to standard releases
The SGLang project says the release brings up to 52% faster decoding with reinforcement-learning sampling masks and up to 12.5× faster ROCm model loading.
TLDR
SGLang announced v0.5.20 with Intel XPU joining its standard releases. The project says reinforcement-learning sampling masks improve rollout reliability and decoding speed, while SGLang-Diffusion gets up to about 38% lower end-to-end latency. Other highlights include SGLang Simulator for scheduler and cache experiments on CPUs, plus new models including GLM-5.3-Flash, Qwen3.8-Flash-Next and K2 Horizon.
