Moonshot AI Releases Kimi K3 MoE Model
Technical report details sparse MoE architecture and efficiency gains over Kimi K2.
TLDR
Moonshot AI published a technical report on its Kimi K3 model along with open weights. The sparse Mixture-of-Experts design routes across hundreds of experts while activating a small subset per token. Architectural updates include Kimi Delta Attention layers, removal of position encodings, and attention residuals that improve retrieval of earlier context. The changes deliver 2.5 times better training efficiency, smaller KV cache footprint at long context, and higher overall scaling performance versus the prior Kimi K2 model. Related open-source components such as MoonEP were also released.
Combined views
214.2K
18 Sources, first seen 67d ago