Report
DeepSeek V4.1-Flash is claimed to cut global KV to 890 bytes per token and shrink its persistent cache by 8x
A post says the model stays competitive on agentic benchmarks and names three techniques behind the claimed cache reductions.
TLDR
A user claims DeepSeek V4.1-Flash cuts global KV to 890 bytes per token and shrinks its persistent cache by 8x while staying competitive on agentic benchmarks. The post names causal encoder-decoder, compressed sparse attention and SWA bounded replay as key techniques, and links a video.
Combined views
23
1 Source, first seen 5h ago
reposts
