DeepSeek Releases V4.1 Flash with Extreme KV Cache Compression for Efficiency
Chinese AI lab DeepSeek launched DeepSeek-V4.1-Flash featuring Causal Encoder-Decoder architecture and FP4 KV caching, compressing the global KV cache to approximately 890 bytes per token. The model supports 1M-token context and is optimized for faster, cheaper inference.
TLDR
Dramatic efficiency gains could significantly reduce API costs and enable longer, more complex agent runs. The release highlights open-weight and Chinese model competition alongside hardware optimization as critical frontiers, potentially reshaping the inference cost landscape.
Combined views
—
1 Source, first seen 11h ago
DeepSeek Releases V4.1 Flash with Extreme KV Cache Compression for Efficiency
Chinese AI lab DeepSeek launched DeepSeek-V4.1-Flash featuring Causal Encoder-Decoder architecture and FP4 KV caching, compressing the global KV cache to approximately 890 bytes per token. The model supports 1M-token context and is optimized for faster, cheaper inference.
TLDR
Dramatic efficiency gains could significantly reduce API costs and enable longer, more complex agent runs. The release highlights open-weight and Chinese model competition alongside hardware optimization as critical frontiers, potentially reshaping the inference cost landscape.