• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    DeepSeek V4.1 Flash’s reported fourth-place sparse-attention benchmark ranking

    A KernelBench-CUDA tester reports runtimes of 0.059–0.736 milliseconds across six test shapes on an RTX PRO 6000 and attributes the entire gap to the top three to unnecessary computation.

    KA
    T(
    LA
    4 Sources, ,

    TLDR

    A user reports that DeepSeek V4.1 Flash placed fourth for DeepSeek Native Sparse Attention on KernelBench-CUDA, behind Fable 5.1, Opus 5 and Fable 5. At an 8K context length, they say its kernel processes 78% of the causal attention blocks even though only about 14% are needed. They attribute the entire performance gap to that excess work. A separate commenter calls V4.1 Flash a major improvement in kernel engineering and says GLM-5.3 Flash is “nowhere close.”

    Combined views

    136.1K

    4 Sources, first seen 19d ago

    Combined views

    136.1K

    4 Sources, first seen 19d ago

    971 likes
    19d ago
    first seen 19d ago
    971 likes
    37 comments
    190 saves
    37 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    37 comments
    190 saves
    37 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    4 Sources

    @elliotarledgeDeepSeek V4.1 Flash on KernelBench-CUDA. DeepSeek Native Sparse Attention for RTX PRO 6000 at 0.50 of the dense-equivalent roofline, fourth on the board. Fable 5.1 is 1.06, Opus 5 is 1.04, Fable 5 is 0.73. The ceiling bills dense attention, so the honest unit is time: 0.059 to 0.736 ms across the six shapes. Inline PTX on SM120: `mma.sync` bf16, `ldmatrix`, xor-swizzled `cp.async`, and an fp32 block-scoring top-8 prologue fused into the attention kernel with the reference's exact tie-break. Then it leaves the sparsity on the table: at 8K context the semantics need about 14% of the causal block triangle and this kernel executes 78% of it. That over-compute is the whole gap to the top three. Rest of this deck: GLM-5.2 Fused MoE: 9.5% of roofline, Opus 5 is 10.7% MegaQwen Decode: 5.4% of roofline, Opus 5 is 6.6% Grid + MinGRU: 29% of roofline, Opus 5 is 196% For this model for DeepSeek, I have not used it much, so I figure I'll at least test it before using it, but it's been a decent general task delegator for now. I haven't really pushed the limits of the model except for this benchmark. https://kernelbench.com/cuda
    @teortaxesTexReally massive improvement with V4.1 Flash on kernel engineering. GLM-5.3 Flash is nowhere close.
    @scaling01RT @elliotarledge: DeepSeek V4.1 Flash on KernelBench-CUDA. DeepSeek Native Sparse Attention for RTX PRO 6000 at 0.50 of the dense-equivale…
    @yacineMTBLove how Claude opus 5 is just straight up wrong. Literal sabotage the AI model

    4 Sources

    @elliotarledgeDeepSeek V4.1 Flash on KernelBench-CUDA. DeepSeek Native Sparse Attention for RTX PRO 6000 at 0.50 of the dense-equivalent roofline, fourth on the board. Fable 5.1 is 1.06, Opus 5 is 1.04, Fable 5 is 0.73. The ceiling bills dense attention, so the honest unit is time: 0.059 to 0.736 ms across the six shapes. Inline PTX on SM120: `mma.sync` bf16, `ldmatrix`, xor-swizzled `cp.async`, and an fp32 block-scoring top-8 prologue fused into the attention kernel with the reference's exact tie-break. Then it leaves the sparsity on the table: at 8K context the semantics need about 14% of the causal block triangle and this kernel executes 78% of it. That over-compute is the whole gap to the top three. Rest of this deck: GLM-5.2 Fused MoE: 9.5% of roofline, Opus 5 is 10.7% MegaQwen Decode: 5.4% of roofline, Opus 5 is 6.6% Grid + MinGRU: 29% of roofline, Opus 5 is 196% For this model for DeepSeek, I have not used it much, so I figure I'll at least test it before using it, but it's been a decent general task delegator for now. I haven't really pushed the limits of the model except for this benchmark. https://kernelbench.com/cuda
    @teortaxesTexReally massive improvement with V4.1 Flash on kernel engineering. GLM-5.3 Flash is nowhere close.
    @scaling01RT @elliotarledge: DeepSeek V4.1 Flash on KernelBench-CUDA. DeepSeek Native Sparse Attention for RTX PRO 6000 at 0.50 of the dense-equivale…
    @yacineMTBLove how Claude opus 5 is just straight up wrong. Literal sabotage the AI model