MiniT2I fine-tune reportedly runs twice as fast with half the tokens
The developer reports the same image quality after switching from a token for every image patch to region tokens.
TLDR
A developer argues that pixel diffusion models waste computation by spending as much on blank sky as on faces. They say their fine-tuned MiniT2I uses region tokens instead, delivering the same quality with half as many tokens at twice the speed. They also report training it at different budgets so it can run with varying amounts of computing power.
Combined views
34
1 Source, first seen 17d ago