• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Reasoning models still need help to control humanoids, one post argues

    Pointing to Google DeepMind and Anthropic’s robotics work, the author sees promise in stronger multimodal reasoning models and presents HumanCLAW as an attempt to glimpse that future.

    DA
    JG
    2 Sources, ,

    TLDR

    One post describes Gemini Robotics 2 as placing an embodied reasoning model above a vision-language-action model for whole-body control, and Anthropic as testing general-purpose reasoning models at several levels of control. The author argues that general-purpose reasoning models cannot yet control a humanoid on their own and need a pretrained control policy underneath. Still, they think a model with strong multimodal and reasoning capabilities could eventually do so, describing HumanCLAW as “our attempt to get an early glimpse of that future.”

    Combined views

    36.8K

    2 Sources, first seen 63d ago

    Combined views

    36.8K

    2 Sources, first seen 63d ago

    187 likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    63d ago
    first seen 63d ago
    187 likes
    11 comments
    139 saves
    38 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    11 comments
    139 saves
    38 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    2 Sources

    @HuggingPapersHumanCLAW: Can Vision-Language Models act through a body? Meta and collaborators introduce a benchmark that decouples action decisions from motor control. Across 1,218 embodied episodes, no VLM solves it — the best reaches only 16.8%. Models lack embodied self-awareness: they lose track of the body they control.
    @KuvviusRobotics felt different this month. Google DeepMind and Anthropic made the same question hard to ignore: can a reasoning model decide what a whole body/robot should do next? Gemini Robotics 2 put an embodied reasoning model above a VLA for whole-body control. Anthropic tested general-purpose reasoning models at several levels of control. - https://deepmind.google/blog/gemini-robotics-2-brings-whole-body-intelligence-to-robots/ - https://anthropic.com/research/claude-plays-robotics The bad news is clear. General-purpose reasoning models still cannot control a humanoid on their own. They need a pretrained policy underneath. Still, I think this might actually work. A strong model (w/ excellent multimodal and reasoning capabilities) could really do it one day. HumanCLAW is our attempt to get an early glimpse of that future.

    2 Sources

    @HuggingPapersHumanCLAW: Can Vision-Language Models act through a body? Meta and collaborators introduce a benchmark that decouples action decisions from motor control. Across 1,218 embodied episodes, no VLM solves it — the best reaches only 16.8%. Models lack embodied self-awareness: they lose track of the body they control.
    @KuvviusRobotics felt different this month. Google DeepMind and Anthropic made the same question hard to ignore: can a reasoning model decide what a whole body/robot should do next? Gemini Robotics 2 put an embodied reasoning model above a VLA for whole-body control. Anthropic tested general-purpose reasoning models at several levels of control. - https://deepmind.google/blog/gemini-robotics-2-brings-whole-body-intelligence-to-robots/ - https://anthropic.com/research/claude-plays-robotics The bad news is clear. General-purpose reasoning models still cannot control a humanoid on their own. They need a pretrained policy underneath. Still, I think this might actually work. A strong model (w/ excellent multimodal and reasoning capabilities) could really do it one day. HumanCLAW is our attempt to get an early glimpse of that future.