Kimi Delta Attention kernel claimed to reach 2.96Γ FlashKDA performance using Kernel Design Agents
The Kernel Design Agents team says it used its system to optimize the Kimi Delta Attention kernel. It released the results and made the kernels open source.
TLDR
The Kernel Design Agents team says its system optimized the Kimi Delta Attention kernel to 2.96Γ FlashKDA performance. It credits iterative work with multiple LLMs, tools for CUDA and other kernel languages, and a self-updating kernel wiki. The team released its results and open-sourced the kernels, while recommending a version verified by FlashInfer for production use.
Combined views
15.7K
2 Sources, first seen 8h ago
