• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    MiniT2I fine-tune runs twice as fast with half the tokens, a user says

    The user says their fine-tuned model uses region tokens instead of treating every image patch as a token in every layer, while maintaining the same quality.

    CL
    EZ
    2 Sources, ,

    TLDR

    A user says pixel diffusion models spend as much compute on blank sky as on faces. Their proposed alternative is a fine-tuned MiniT2I that uses region tokens, which they claim delivers the same quality with half as many tokens at twice the speed. They also say it was trained at different compute budgets to support flexible inference.

    Combined views

    67.9K

    2 Sources, first seen 19d ago

    Combined views

    67.9K

    2 Sources, first seen 19d ago

    588 likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    19d ago
    first seen 19d ago
    588 likes
    16 comments
    413 saves
    48 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    16 comments
    413 saves
    48 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    2 Sources

    @eduard__zamfirPixel diffusion models spend as much compute on blank sky as on faces. Every patch is a token in every layer. Our fine-tuned MiniT2I runs on region tokens instead. Half as many tokens, same quality, twice as fast, trained at different budgets for elastic inference.
    @_chenglouQuad trees are the BPE of vision

    2 Sources

    @eduard__zamfirPixel diffusion models spend as much compute on blank sky as on faces. Every patch is a token in every layer. Our fine-tuned MiniT2I runs on region tokens instead. Half as many tokens, same quality, twice as fast, trained at different budgets for elastic inference.
    @_chenglouQuad trees are the BPE of vision