Reasoning models still need help to control humanoids, one post argues
Pointing to Google DeepMind and Anthropic’s robotics work, the author sees promise in stronger multimodal reasoning models and presents HumanCLAW as an attempt to glimpse that future.
TLDR
One post describes Gemini Robotics 2 as placing an embodied reasoning model above a vision-language-action model for whole-body control, and Anthropic as testing general-purpose reasoning models at several levels of control. The author argues that general-purpose reasoning models cannot yet control a humanoid on their own and need a pretrained control policy underneath. Still, they think a model with strong multimodal and reasoning capabilities could eventually do so, describing HumanCLAW as “our attempt to get an early glimpse of that future.”
Combined views
36.8K
2 Sources, first seen 63d ago