• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    “Robot-use agents” mark an important robotics shift, a writer argues

    A reply says language models could improve with a closed feedback loop, accurate task checks and safety guardrails—but argues the system needed for that setup is still far off.

    CM
    PI
    DP
    26 Sources, ,

    TLDR

    Pointing to demos of AI agents such as Claude controlling robots, one writer calls “robot-use agents” an important change for robotics. The linked blog describes the idea as a language model using a robot as a tool, much as it uses a calculator or web search.

    A reply offers conditional optimism: With a closed-loop system that feeds results back into the process, accurate task verification and safety guardrails, language models could improve iteratively. But the commenter argues that the technical stack needed for that setup is still far off.

    Combined views

    370.8K

    26 Sources, first seen 23d ago

    Combined views

    370.8K

    26 Sources, first seen 23d ago

    2.6K likes
    23d ago
    first seen 23d ago
    2.6K likes
    116 comments
    1.3K saves
    385 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    116 comments
    1.3K saves
    385 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    26 Sources

    @phillip_isolaRecently, there have been a lot of impressive demos of AI agents, like Claude, controlling robots. I wrote a short blog post with my thoughts on the advent of these "robot-use agents." https://web.mit.edu/phillipi/www/writing/robot-use-agents.html I think it's an important change in the trajectory of robotics!
    @georgiagkioxari@phillip_isola If there is a closed-loop system, with accurate task verification and safety-guardrails, then I too am confident that LLMs can hillclimb. But I do think that the stack needed to get to this setup is far from where we are today.
    @vincesitzmannI agree with Phil's take here: The progress of LLMs on controlling robots is quite interesting. Intuitively, this makes sense: controlling a robot is not so different from computer use, and an agent that is good at computer use is probably also good at controlling a robot and vice versa - credit to my student @RyuHyunwoooo to pointing that out to me originally!
    @ManlingLi_RT @Kuvvius: Robotics felt different this month. Google DeepMind and Anthropic made the same question hard to ignore: can a reasoning model…
    @IMordatch@phillip_isola I think it's not certain yet how directly/frequently involved the models will ends up being (as opposed to creating something to delegate to), but I'm really glad we built up the audacity and fearlessness to ask for a use case like this!
    @chooi_jeqToday from Phillip Isola, MIT professor and computer vision legend: “Every robot with an internet connection becomes a potential tool for Fable, Astra, and other AIs.” 🧵
    @Ken_GoldbergIt feels like an inflection point — one issue is whether LLM‘s role is online or offline. The latter can be more useful for industry, but both make sense.
    @DimitrisPapailEmbodied experience as a tool or perhaps the robotic body as a physical world harness. Very interesting thought experiment that’s already happening (ref Claude with a robotic arm) and likely to proliferate since you can further perform RL on such an agent on actual physical tasks. I feel we’re at the equivalent level of pre GSM8k but soon will see rapid progress
    @JitendraMalikCVI disagree. You need high control frequency controllers/policies that deal with forces and contacts. If you believe what you said, here is a simple challenge: implement policies for quadruped locomotion e.g.(https://vision-locomotion.github.io/ RSS 2021, CoRL 2022) in your favorite LLM and test them on different terrains. Doing a quasi-static manipulation task with parallel jaw grippers does not equate to dexterity (You may want to talk to your colleague Pulkit!). High level planning can certainly be done by LLMs but that is not where the hard part is.
    @chrmanningBy somewhat weakening the claims at the start of Phillip’s piece, can’t you both be right? Yes, you also need low-level robot control software, but the high-level LLM can use it as a tool, and most of the picture Phillip paints can still be true whereby robot intelligence quickly takes off via using LLMs for higher-level planning and control, roughly corresponding to conscious human planning and control.

    26 Sources

    @phillip_isolaRecently, there have been a lot of impressive demos of AI agents, like Claude, controlling robots. I wrote a short blog post with my thoughts on the advent of these "robot-use agents." https://web.mit.edu/phillipi/www/writing/robot-use-agents.html I think it's an important change in the trajectory of robotics!
    @georgiagkioxari@phillip_isola If there is a closed-loop system, with accurate task verification and safety-guardrails, then I too am confident that LLMs can hillclimb. But I do think that the stack needed to get to this setup is far from where we are today.
    @vincesitzmannI agree with Phil's take here: The progress of LLMs on controlling robots is quite interesting. Intuitively, this makes sense: controlling a robot is not so different from computer use, and an agent that is good at computer use is probably also good at controlling a robot and vice versa - credit to my student @RyuHyunwoooo to pointing that out to me originally!
    @ManlingLi_RT @Kuvvius: Robotics felt different this month. Google DeepMind and Anthropic made the same question hard to ignore: can a reasoning model…
    @IMordatch@phillip_isola I think it's not certain yet how directly/frequently involved the models will ends up being (as opposed to creating something to delegate to), but I'm really glad we built up the audacity and fearlessness to ask for a use case like this!
    @chooi_jeqToday from Phillip Isola, MIT professor and computer vision legend: “Every robot with an internet connection becomes a potential tool for Fable, Astra, and other AIs.” 🧵
    @Ken_GoldbergIt feels like an inflection point — one issue is whether LLM‘s role is online or offline. The latter can be more useful for industry, but both make sense.
    @DimitrisPapailEmbodied experience as a tool or perhaps the robotic body as a physical world harness. Very interesting thought experiment that’s already happening (ref Claude with a robotic arm) and likely to proliferate since you can further perform RL on such an agent on actual physical tasks. I feel we’re at the equivalent level of pre GSM8k but soon will see rapid progress
    @JitendraMalikCVI disagree. You need high control frequency controllers/policies that deal with forces and contacts. If you believe what you said, here is a simple challenge: implement policies for quadruped locomotion e.g.(https://vision-locomotion.github.io/ RSS 2021, CoRL 2022) in your favorite LLM and test them on different terrains. Doing a quasi-static manipulation task with parallel jaw grippers does not equate to dexterity (You may want to talk to your colleague Pulkit!). High level planning can certainly be done by LLMs but that is not where the hard part is.
    @chrmanningBy somewhat weakening the claims at the start of Phillip’s piece, can’t you both be right? Yes, you also need low-level robot control software, but the high-level LLM can use it as a tool, and most of the picture Phillip paints can still be true whereby robot intelligence quickly takes off via using LLMs for higher-level planning and control, roughly corresponding to conscious human planning and control.