Yuxiao Qu Announces R3 Robot Reasoning via Reinforcement Learning
Post highlights new approach to scaling test-time compute in robotic perception-action loops.
TLDR
PhD researcher Yuxiao Qu posted an announcement of new work called R3. The title given is Training Robots to Reason in Natural Language via Reinforcement Learning. Qu states that scaling test-time compute for robots is not simply a matter of using a stronger reasoning LLM or VLM. Instead the extra computation occurs inside a closed perception-action loop. In that setting intermediate reasoning must connect directly to ongoing actions. The post lists Qu's current role at CMU and earlier affiliations at UW-Madison, UW, and CUHK.
Combined views
7.7K
5 Sources, first seen 29d ago