LiFT loops part of a diffusion transformer to scale inference depth
The project team reports lower ImageNet FID than dense DiT-XL/2 while using fewer parameters and FLOPs.
TLDR
LiFT makes part of a diffusion transformer recurrent so it can spend more compute inside each sampling step. The team reports a 3.34 lower ImageNet FID than dense DiT-XL/2, with 60% fewer parameters, 32% fewer training FLOPs and 52% fewer inference FLOPs. A supervisor says it can loop beyond its training depth at inference without retraining, early exits or other changes.
Combined views
9.9K
4 Sources, first seen ago
