MiniT2I fine-tune runs twice as fast with half the tokens, a user says
The user says their fine-tuned model uses region tokens instead of treating every image patch as a token in every layer, while maintaining the same quality.
TLDR
A user says pixel diffusion models spend as much compute on blank sky as on faces. Their proposed alternative is a fine-tuned MiniT2I that uses region tokens, which they claim delivers the same quality with half as many tokens at twice the speed. They also say it was trained at different compute budgets to support flexible inference.
Combined views
67.9K
2 Sources, first seen 19d ago