• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Announcement

    BIND links candidate robot actions to 2D image features

    A post introducing the action head says it makes robot policies more data-efficient and robust to unfamiliar camera views and object positions.

    CS
    1 Source, 6h ago, first seen 6h ago

    TLDR

    A post introducing BIND argues that robot image encoders capture geometry and meaning, while policies built on them can need hundreds of demonstrations for simple pick-and-place tasks and break after small camera shifts. BIND links each candidate action to the image feature where the robot’s end effector would move. The post claims this makes policies more data-efficient and robust to unfamiliar viewpoints and object positions, using RGB input without 3D sensors.

    Combined views

    4.7K

    1 Source, first seen 6h ago

    Combined views

    4.7K

    1 Source, first seen 6h ago

    93 likes
    93 likes
    5 comments
    96 saves
    11 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Featured Source
    5 comments
    96 saves
    11 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @omcamsmithA paradox in robot learning: robot policies are spatially dumb and data-inefficient but their image features are spatially robust and semantically rich. Image encoders are spatially smart — features embed semantics + geometry + are multiview-consistent ... but robot policies we build upon them are spatially dumb — hundreds of demos to learn simple pick+place and a small camera bump breaks them To better bridge this gap, check out BIND: a new action head that binds each candidate robot action to its projecting 2D image feature(s) BIND basically offers the network the info of '‘choosing this candidate robot action would move the robot EEF to this image feature." Result: policies much more data efficient and robust to OOD viewpoints and OOD object positions (same rgb input, no 3D sensors, dense EEF trajectories out) 🧵6h