DeepSeek V4.1 Flash’s reported fourth-place sparse-attention benchmark ranking
A KernelBench-CUDA tester reports runtimes of 0.059–0.736 milliseconds across six test shapes on an RTX PRO 6000 and attributes the entire gap to the top three to unnecessary computation.
TLDR
A user reports that DeepSeek V4.1 Flash placed fourth for DeepSeek Native Sparse Attention on KernelBench-CUDA, behind Fable 5.1, Opus 5 and Fable 5. At an 8K context length, they say its kernel processes 78% of the causal attention blocks even though only about 14% are needed. They attribute the entire performance gap to that excess work. A separate commenter calls V4.1 Flash a major improvement in kernel engineering and says GLM-5.3 Flash is “nowhere close.”
Combined views
136.1K
4 Sources, first seen 19d ago