• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Yuxiao Qu Announces R3 Robot Reasoning via Reinforcement Learning

    Post highlights new approach to scaling test-time compute in robotic perception-action loops.

    AK
    TW
    YQ
    5 Sources, 29d ago, first seen 29d ago

    TLDR

    PhD researcher Yuxiao Qu posted an announcement of new work called R3. The title given is Training Robots to Reason in Natural Language via Reinforcement Learning. Qu states that scaling test-time compute for robots is not simply a matter of using a stronger reasoning LLM or VLM. Instead the extra computation occurs inside a closed perception-action loop. In that setting intermediate reasoning must connect directly to ongoing actions. The post lists Qu's current role at CMU and earlier affiliations at UW-Madison, UW, and CUHK.

    Combined views

    7.7K

    5 Sources, first seen 29d ago

    Combined views

    7.7K

    5 Sources, first seen 29d ago

    84 likes
    84 likes
    5 comments
    49 saves
    14 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    5 comments
    49 saves
    14 reposts
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    5 Sources

    @QuYuxiaoExcited to share a new work! 🚨 R3: Training Robots to Reason in Natural Language via Reinforcement Learning 🔥 Scaling test-time compute for robots is not just giving them a stronger reasoning LLM/VLM. In robotics, extra test-time computation happens inside a closed perception–action loop, where intermediate reasoning can change what the robot observes and does next.
    @tw_killianRT @QuYuxiao: Excited to share a new work! 🚨 R3: Training Robots to Reason in Natural Language via Reinforcement Learning 🔥 Scaling test-…
    @aviral_kumar2In our new work ⬇️, we develop robot VLMs that can improve their own reasoning simply by practicing to *explain* demo data. Trained robot VLMs then control low-level VLAs. This is important as reasoning in language can provide a cheap axis for scaling TTC even for robots and hence, enable fast autonomous improvement & generalization in new scenarios. Many SOTA robot learning methods still handcraft reasoning annotations (or equivalents). But this is suboptimal if we take a lesson from how LLMs are trained to reason.

    5 Sources

    @QuYuxiaoExcited to share a new work! 🚨 R3: Training Robots to Reason in Natural Language via Reinforcement Learning 🔥 Scaling test-time compute for robots is not just giving them a stronger reasoning LLM/VLM. In robotics, extra test-time computation happens inside a closed perception–action loop, where intermediate reasoning can change what the robot observes and does next.
    @tw_killianRT @QuYuxiao: Excited to share a new work! 🚨 R3: Training Robots to Reason in Natural Language via Reinforcement Learning 🔥 Scaling test-…
    @aviral_kumar2In our new work ⬇️, we develop robot VLMs that can improve their own reasoning simply by practicing to *explain* demo data. Trained robot VLMs then control low-level VLAs. This is important as reasoning in language can provide a cheap axis for scaling TTC even for robots and hence, enable fast autonomous improvement & generalization in new scenarios. Many SOTA robot learning methods still handcraft reasoning annotations (or equivalents). But this is suboptimal if we take a lesson from how LLMs are trained to reason.