• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Announcement

    Uni-LaDiR proposes a shared reasoning space for images, text and 3D point clouds

    The introductory post says the framework generates “latent thoughts” with diffusion for VLMs and VLAs.

    Lianhui Qin@COLM2026LQ
    Murray KangMK
    2 Sources, ,

    TLDR

    A post introducing Uni-LaDiR argues that reasoning should happen in an abstract latent space rather than separately in images, text and 3D point clouds. It describes a diffusion-based framework that generates “latent thoughts” for vision-language models (VLMs) and vision-language-action models (VLAs).

    Combined views

    1.3K

    2 Sources, first seen 2h ago

    Combined views

    1.3K

    2 Sources, first seen 2h ago

    18 likes
    2h ago
    first seen 2h ago
    18 likes
    3 comments
    5 saves
    5 reposts
    3 comments
    5 saves
    5 reposts
    Featured Source

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    Murray Kang@haoqik322Can different modalities be unified into one space for reasoning? Humans don’t reason separately in images, text, and 3D point clouds — we don’t “reason in pixels,” then switch to “reasoning in words.” We believe thinking should live in a more abstract latent space, independent of whether the CoT data comes from images, text, or 3D point clouds. Introducing Uni-LaDiR (Unified Latent Diffusion Reasoner), a unified reasoning framework for both VLMs and VLAs that generates latent thoughts with diffusion.2h
    Lianhui Qin@COLM2026@LianhuiqUnifying multimodal reasoning in latent space could let physical AI combine and reuse world context more efficiently across tasks. Consider a factory cart: navigation needs its location, manipulation needs its geometry, and inspection needs images of what has changed. Same object. Different perspectives. Each contributes information the others may miss. Uni-LaDiR takes a step toward this vision, learning a common latent reasoning interface from text, images, 3D geometry, and robot state. The opportunity: bring complementary information into a shared reasoning space, so each task can draw on the world context it needs.2h

    2 Sources

    Murray Kang@haoqik322Can different modalities be unified into one space for reasoning? Humans don’t reason separately in images, text, and 3D point clouds — we don’t “reason in pixels,” then switch to “reasoning in words.” We believe thinking should live in a more abstract latent space, independent of whether the CoT data comes from images, text, or 3D point clouds. Introducing Uni-LaDiR (Unified Latent Diffusion Reasoner), a unified reasoning framework for both VLMs and VLAs that generates latent thoughts with diffusion.2h
    Lianhui Qin@COLM2026@LianhuiqUnifying multimodal reasoning in latent space could let physical AI combine and reuse world context more efficiently across tasks. Consider a factory cart: navigation needs its location, manipulation needs its geometry, and inspection needs images of what has changed. Same object. Different perspectives. Each contributes information the others may miss. Uni-LaDiR takes a step toward this vision, learning a common latent reasoning interface from text, images, 3D geometry, and robot state. The opportunity: bring complementary information into a shared reasoning space, so each task can draw on the world context it needs.2h