Announcement
KV caching cuts latency in image-generation flow-model benchmarks, but uses more GPU memory
Sayak Paul reports 23%–28.5% lower latency with one reference image in Flux.2 Klein tests, alongside 2.1% higher peak GPU memory.
TLDR
Sayak Paul explains how caching key and value tensors for reference-image tokens lets flow-based image generators reuse them across denoising steps. In his Flux.2 Klein benchmarks, caching reduced latency by 23%–28.5% with one reference image and 42.1%–51.3% with three; peak GPU memory rose 2.1% and 11.7%, respectively. He also reports a 44.7% latency reduction with QwenImage 2.1, with 5.7% higher peak GPU memory.
Combined views
15K
2 Sources, first seen ago