Announcement
Uni-LaDiR proposes a shared reasoning space for images, text and 3D point clouds
The introductory post says the framework generates “latent thoughts” with diffusion for VLMs and VLAs.
TLDR
A post introducing Uni-LaDiR argues that reasoning should happen in an abstract latent space rather than separately in images, text and 3D point clouds. It describes a diffusion-based framework that generates “latent thoughts” for vision-language models (VLMs) and vision-language-action models (VLAs).
Combined views
1.3K
2 Sources, first seen ago
