Report
LDM-is-AE proposes image generation without a pretrained VAE
A post sharing the paper says it reinterprets DiT as an autoencoder and learns latent representations jointly with denoising.
TLDR
The post says LDM-is-AE’s diffusion-native latent removes the need for a pretrained VAE and improves ImageNet image generation to an FID of 1.80.
Combined views
3
1 Source, first seen 11h ago
9 reposts