• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    LIT’s approach to teaching goal-directed motion before adding vision

    The post argues that learning movement without images can avoid visual shortcuts. Later vision training is meant to support the action expert’s existing spatial understanding.

    CP
    KK
    JD
    3 Sources, ,

    TLDR

    A post explaining LIT describes two training stages. First, an “action expert” learns to move toward a spatial goal without images. Then, vision is introduced through a latent interface trained to recover that same goal. The author’s rationale is to establish goal-directed movement first, then encourage visual information to support what matters for the action.

    Combined views

    15.4K

    3 Sources, first seen 20d ago

    Combined views

    15.4K

    3 Sources, first seen 20d ago

    204 likes
    20d ago
    first seen 20d ago
    204 likes
    7 comments
    168 saves
    41 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    7 comments
    168 saves
    41 reposts
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    3 Sources

    @DJiafeiAre you training your VLAs or WAMs correctly? Your action expert may be learning vision–action shortcuts that undermine generalization beyond the training distribution. Introducing Latent Interface Training (LIT 🔥): Learn to act first, then learn how to use vision. 🧵👇
    @kastnerkyleRT @DJiafei: 3/🧵 Why does LIT make intuitive sense? First, give the action expert a spatial goal and teach it how to move there without im…
    @chris_j_paxtonRT @DJiafei: Are you training your VLAs or WAMs correctly? Your action expert may be learning vision–action shortcuts that undermine gener…

    3 Sources

    @DJiafeiAre you training your VLAs or WAMs correctly? Your action expert may be learning vision–action shortcuts that undermine generalization beyond the training distribution. Introducing Latent Interface Training (LIT 🔥): Learn to act first, then learn how to use vision. 🧵👇
    @kastnerkyleRT @DJiafei: 3/🧵 Why does LIT make intuitive sense? First, give the action expert a spatial goal and teach it how to move there without im…
    @chris_j_paxtonRT @DJiafei: Are you training your VLAs or WAMs correctly? Your action expert may be learning vision–action shortcuts that undermine gener…