• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Reaction

    Multimodal LLM is claimed to tackle a robotics task it was never deliberately trained for

    The demonstrator says DSV4.1F uses unfamiliar cameras and motors; the video runs at 10x speed.

    Chris PaxtonCP
    Harrison KinsleyHK
    2 Sources, ,

    TLDR

    A demonstrator says DSV4.1F, a multimodal LLM, tackles a task it was never deliberately trained for using cameras it has never seen and motors it did not learn on. The video runs at 10x speed because the system is currently slow. The demonstrator argues that general-purpose models may displace much robotics task training, while reinforcement learning remains valuable for low-level skills like gait and reaching.

    Combined views

    4.4K

    2 Sources, first seen 3h ago

    Combined views

    4.4K

    2 Sources, first seen 3h ago

    74 likes
    3h ago
    first seen 3h ago
    74 likes
    12 comments
    22 saves
    12 reposts
    12 comments
    22 saves
    12 reposts
    Featured Source

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    Harrison Kinsley@SentdexNo training, no sim2real gaps, just a prompt and a dream. The problem in robotics today is every task is a complex combination of locomotion, manipulation, vision, planning, and a bunch of other things that fall apart once you diverge from training data or characteristics from sim. This is why you see so many demos with perfect lighting conditions and always the same, or precisely placed objects, specifically chosen for the task and demo. You never see these things deployed in the real world despite seeing demos that look convincing, because everything falls flat in the real world with real unseen conditions. Every single deployment is a unique edgecase. Here, it's just a multimodal LLM (DSV4.1F) doing a task it was never deliberately trained for, using cameras it's never seen, servos/motors it didn't learn on, in an environment it's never dealt with and a form factor never tested. It is currently slow, this video is 10x speed, but I am very confident it can be sped up as is, but also this is basically the slowest this level of intelligence will ever be again. The more I work with general purpose multimodal LLMs, the more I am realizing that ACT, VLAs, and all this other stuff is probably just a sideshow and a distraction, and it's just going to largely get eaten up by general purpose models that just happen to also do robotics. I don't think it'll always be LLMs, but the fact that LLMs can do some of the most challenging robotics tasks tells me it's kind of over for a large class of robotics training. This does not include RL, however, for low level things, like gait, reaching, and other capabilities or "skills" for hardware. RL is still a king here and you really do want these models to be super fast and super small. RL is very good here and safe I'd say. Decision models like Clef, Jev...etc remain mostly useless imho for robotics. I look forward to someone proving out a useful recipe that works these things into robotics, but I don't currently see it.3h
    Chris Paxton@chris_j_paxtonRT @Sentdex: No training, no sim2real gaps, just a prompt and a dream. The problem in robotics today is every task is a complex combinatio…2h

    2 Sources

    Harrison Kinsley@SentdexNo training, no sim2real gaps, just a prompt and a dream. The problem in robotics today is every task is a complex combination of locomotion, manipulation, vision, planning, and a bunch of other things that fall apart once you diverge from training data or characteristics from sim. This is why you see so many demos with perfect lighting conditions and always the same, or precisely placed objects, specifically chosen for the task and demo. You never see these things deployed in the real world despite seeing demos that look convincing, because everything falls flat in the real world with real unseen conditions. Every single deployment is a unique edgecase. Here, it's just a multimodal LLM (DSV4.1F) doing a task it was never deliberately trained for, using cameras it's never seen, servos/motors it didn't learn on, in an environment it's never dealt with and a form factor never tested. It is currently slow, this video is 10x speed, but I am very confident it can be sped up as is, but also this is basically the slowest this level of intelligence will ever be again. The more I work with general purpose multimodal LLMs, the more I am realizing that ACT, VLAs, and all this other stuff is probably just a sideshow and a distraction, and it's just going to largely get eaten up by general purpose models that just happen to also do robotics. I don't think it'll always be LLMs, but the fact that LLMs can do some of the most challenging robotics tasks tells me it's kind of over for a large class of robotics training. This does not include RL, however, for low level things, like gait, reaching, and other capabilities or "skills" for hardware. RL is still a king here and you really do want these models to be super fast and super small. RL is very good here and safe I'd say. Decision models like Clef, Jev...etc remain mostly useless imho for robotics. I look forward to someone proving out a useful recipe that works these things into robotics, but I don't currently see it.3h
    Chris Paxton@chris_j_paxtonRT @Sentdex: No training, no sim2real gaps, just a prompt and a dream. The problem in robotics today is every task is a complex combinatio…2h