Bidirectional Diffusion Bridges Proposed for Multimodality Translation
Paper proposes BIT method that starts from text and interpolates directly into images.
TLDR
The paper "There and Back Again: Bidirectional Diffusion Bridges for Multimodality Translation" proposes BIT, a method that starts from text and interpolates directly into images instead of noise. According to the quoted post, this supplies a source-aware generative path for more diverse sampling and allows the process to run backward from image to text. The work links to the arXiv entry at arxiv.org/abs/2608.27885, a project site at bit-diffusion.github.io, and a GitHub repository under the gabeguo account.
Combined views
45K
4 Sources, first seen 31d ago
