• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Four-legged robot moves chairs with contact-seeking training, a post reports

    The method guides the robot toward possible contact points, then gradually reduces that guidance during training so it can focus on completing the task, according to the post.

    1 Source, 23d ago, first seen 23d ago

    TLDR

    The post describes a reinforcement-learning problem: when task rewards stay at zero until contact, a robot can get stuck optimizing smoothness and energy penalties instead. It credits researchers from the University of Pisa, ETH Zürich and Nvidia with adding a separate “critic,” or evaluator, trained on contact-seeking rewards. Its influence decreases during training, shifting emphasis toward task performance. According to the post, candidate contact points come from a general-purpose grasping algorithm. It reports tests on a real four-legged robot with a manipulator that moved chairs, transferred to unseen IKEA furniture without additional training, and recovered from failed contact attempts.

    Combined views

    —

    1 Source, first seen 23d ago

    Combined views

    —

    1 Source, first seen 23d ago

    — likes
    — likes
    — comments
    — saves
    — reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    — comments
    — saves
    — reposts

    1 Source

    @stepjamUKTeaching a quadruped robot to push an object it can't grasp is a hard RL exploration problem. The task reward stays at zero until contact happens, so a standard single-critic PPO ends up optimizing smoothness and energy penalties instead, and never finds contact at all. Members from @Unipisa, ETH Zürich & @nvidia address this with a separate exploration critic trained on a dense contact-seeking reward, guiding the end-effector toward candidate contact points, then decaying its weight over training so the policy shifts from contact-seeking to task-optimal once the physics has actually been found. Candidate points come from a general-purpose grasping algorithm, so the approach generalizes across object geometries without hand-tuning per task. Validated on a real quadrupedal mobile manipulator transporting chairs, transferring zero-shot to unseen IKEA furniture, recovering from failed contact attempts, and remaining stable under loads beyond the robot's rated capacity. A clean approach to a common problem: rather than hand-shaping one messy scalar reward, split exploration from task performance and let one decay into the other. Paper: https://tolomeis.github.io/contact-guided-exp/assets/contact_guided_exp_RAL.pdf Project page: https://tolomeis.github.io/contact-guided-exp/

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    1 Source

    @stepjamUKTeaching a quadruped robot to push an object it can't grasp is a hard RL exploration problem. The task reward stays at zero until contact happens, so a standard single-critic PPO ends up optimizing smoothness and energy penalties instead, and never finds contact at all. Members from @Unipisa, ETH Zürich & @nvidia address this with a separate exploration critic trained on a dense contact-seeking reward, guiding the end-effector toward candidate contact points, then decaying its weight over training so the policy shifts from contact-seeking to task-optimal once the physics has actually been found. Candidate points come from a general-purpose grasping algorithm, so the approach generalizes across object geometries without hand-tuning per task. Validated on a real quadrupedal mobile manipulator transporting chairs, transferring zero-shot to unseen IKEA furniture, recovering from failed contact attempts, and remaining stable under loads beyond the robot's rated capacity. A clean approach to a common problem: rather than hand-shaping one messy scalar reward, split exploration from task performance and let one decay into the other. Paper: https://tolomeis.github.io/contact-guided-exp/assets/contact_guided_exp_RAL.pdf Project page: https://tolomeis.github.io/contact-guided-exp/
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet