• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Announcement

    LiFT loops part of a diffusion transformer to scale inference depth

    The project team reports lower ImageNet FID than dense DiT-XL/2 while using fewer parameters and FLOPs.

    Kosta DerpanisKD
    PedroPE
    Cees SnoekCS
    4 Sources, ,

    TLDR

    LiFT makes part of a diffusion transformer recurrent so it can spend more compute inside each sampling step. The team reports a 3.34 lower ImageNet FID than dense DiT-XL/2, with 60% fewer parameters, 32% fewer training FLOPs and 52% fewer inference FLOPs. A supervisor says it can loop beyond its training depth at inference without retraining, early exits or other changes.

    Combined views

    9.9K

    4 Sources, first seen 4h ago

    Combined views

    9.9K

    4 Sources, first seen 4h ago

    171 likes

    Useful links

    arXiv.org

    LiFT: Loop Flow Transformers

    arXiv.org

    Looped Diffusion Transformer
    4h ago
    first seen 4h ago
    171 likes
    8 comments
    127 saves
    50 reposts
    8 comments
    127 saves
    50 reposts

    The team behind LiFT, short for Loop Flow Transformers, has introduced a flow-matching transformer that makes part of a diffusion transformer recurrent. Project co-lead Pedro Curvo says the design lets the model spend more compute inside each sampling step instead of scaling inference only by adding sampling steps.

    Featured Source

    That recurrent core is intended to let LiFT scale its depth at inference with a relatively small architectural change, Curvo wrote.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Useful Links

    arXiv.org

    LiFT: Loop Flow Transformers

    arXiv.org

    Looped Diffusion Transformer
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    The team's ImageNet claims

    On ImageNet, the team reports that LiFT achieved a 3.34 lower FID than dense DiT-XL/2 while using 60% fewer parameters, 32% fewer training FLOPs and 52% fewer inference FLOPs. The public posts announcing LiFT do not spell out the benchmark methodology, so the figures should be read as team-reported results.

    Project supervisor Cees Snoek says LiFT can loop at inference far beyond its training depth without retraining, early exits or other modifications.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    4 Sources

    Pedro@pmpcurvo♾️ LiFT: Loop Flow Transformers 🧵👇 Diffusion models usually spend more inference compute by taking more sampling steps. What if we could also spend compute inside each step? LiFT turns part of a DiT into a recurrent core, letting us scale depth at inference with only a small architectural change. On ImageNet, it beats dense DiT-XL/2 with 3.34 lower FID, while using 60% fewer parameters, 32% fewer training FLOPs, and 52% fewer inference FLOPs. Jointly led with Mohammad Mahdi Derakhshani 👑(@mmderakhshani) And amazing supervisors Cees Snoek (@cgmsnoek), Jan-Willem van de Meent (@jwvdm) and Gertjan J. Burghouts (@gjburghouts)4h
    Cees Snoek@cgmsnoek🛗 Introducing LiFT: A flow matching transformer that loops at inference far beyond its training depth, with no retraining, early exits, or other modifications. 💡3h
    Kosta Derpanis@CSProfKGDRT @cgmsnoek: 🛗 Introducing LiFT: A flow matching transformer that loops at inference far beyond its training depth, with no retraining, ea…2h

    Useful Links

    arXiv.org

    LiFT: Loop Flow Transformers

    arXiv.org

    Looped Diffusion Transformer

    4 Sources

    Pedro@pmpcurvo♾️ LiFT: Loop Flow Transformers 🧵👇 Diffusion models usually spend more inference compute by taking more sampling steps. What if we could also spend compute inside each step? LiFT turns part of a DiT into a recurrent core, letting us scale depth at inference with only a small architectural change. On ImageNet, it beats dense DiT-XL/2 with 3.34 lower FID, while using 60% fewer parameters, 32% fewer training FLOPs, and 52% fewer inference FLOPs. Jointly led with Mohammad Mahdi Derakhshani 👑(@mmderakhshani) And amazing supervisors Cees Snoek (@cgmsnoek), Jan-Willem van de Meent (@jwvdm) and Gertjan J. Burghouts (@gjburghouts)4h
    Cees Snoek@cgmsnoek🛗 Introducing LiFT: A flow matching transformer that loops at inference far beyond its training depth, with no retraining, early exits, or other modifications. 💡3h
    Kosta Derpanis@CSProfKGDRT @cgmsnoek: 🛗 Introducing LiFT: A flow matching transformer that loops at inference far beyond its training depth, with no retraining, ea…2h