Luke Huang Posts First Blackwell GEMM Kernel Blog
Post covers GPU architecture, CuTeDSL, loop tiling with TMA and TMEM, plus swizzling.
TLDR
Luke D. Huang, a MIT student in CS and physics who previously worked at Databricks and AWS, posted on X to announce his first blog. The entry is part 1 of a series on writing high-performance GEMM kernels for Blackwell GPUs. It addresses basic GPU architecture, CuTeDSL fundamentals, loop tiling using TMA and TMEM, and swizzling. The post is presented as the start of a longer effort to reach speed-of-light kernel performance.
Combined views
38.3K
1 Source, first seen 26d ago