• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
Technology

Rhoda AI Study: Web Videos Boost Robot Factory Performance

Study tests if larger models trained on web videos improve factory manipulation performance.

TM
AL
JS
16 Sources, 21d ago, first seen 21d ago

TLDR

Engineers at Rhoda AI trained video models on general web clips at increasing scales of model size and compute. They then evaluated the resulting policies on a single real-world industrial manipulation job. The team observed that better prediction of held-out web video correlated with stronger post-training robot performance. Tongzhou Mu led the work, including all training and evaluations. The findings appear in a new blog post from the company.

Combined views

159.4K

16 Sources, first seen 21d ago

717 likes75 comments259 saves74 reposts
Featured Source

Combined views

159.4K

16 Sources, first seen 21d ago

717 likes75 comments259 saves74 reposts
Seven pre-trained video-model checkpoints on one chart. Task performance on a real robot (at-speed completion rate: every requirement met within 100 s, 0 to 100%) rises with pre-training quality (DINO FD on held-out web video, log scale, axis reversed so better is to the right). XS, S, M, L are model sizes. The four M points are one model at 0.08×, 0.18×, 0.37× and 1× of its pre-training compute.

Sentiment

Positive94.1%5.9%Negative

Summary

Many accounts praised the robotics research for confirming that scaling web video pre-training improves real-world robot performance on complex industrial tasks.

Based on 34 sentiment-bearing replies from 34 accounts across 4 conversations.

Sentiment

Positive94.1%5.9%Negative

Summary

Many accounts praised the robotics research for confirming that scaling web video pre-training improves real-world robot performance on complex industrial tasks.

Based on 34 sentiment-bearing replies from 34 accounts across 4 conversations.

16 Sources

@tongzhou_mu1/ Setup. The video model is pre-trained on general web video, which contains no actions, then post-trained on robot demos of one industrial task, bearing unpacking. Each policy gets 100 to a few hundred robot trials, over 200 robot hours in total, every trial set up to a written procedure and audited afterwards. Evaluating on a real robot is hard, so the post describes how we do it in detail.
@aliteracytongzhou led this entire work from start to finish (including training every model, finding bugs, and supervising every evaluation trial himself personally) - an absolute 🐐!! what excited me most about these results is that we can now use DINO FD as a predictor for post-training performance all the way from pre-training
@startupjag@tongzhou_mu Scaling the AI model is one thing. Understanding whether it actually translates to better real-world robot performance is another. Fantastic work from @Tongzhou Mu and @RhodaAI putting it to the test.
@RhodaAIToday marks week one of our new blog series. First one's from @tongzhou_mu . Will scaling web-video pre-training lead to better robot policies?
@adam_patniRigorous evaluation on real world tasks is what will enable the best companies to outperform the rest. Scaling model size, compute, and improving web-data generation quality have empirically produced better robot policies across hundreds/thousands of trials at Rhoda. Major props to Tongzhou, who literally froze every single non-model thing in our office to produce this. I was literally maintaining months old infra to keep this eval going.
@felixwyw@tongzhou_mu Congrats @tongzhou_mu amazing work!
@GordonWetzstein@tongzhou_mu It's great to see a principled study on how pretraining at scale helps real-world robotics manipulation tasks. The insight is intuitive: pre-training on large web-video datasets helps downstream robotics tasks a lot!
@charles_rqiGreat to see a principled study of video model pretraining for robot manipulation. I especially like the evaluation setup: a real-world industrial task and a strong, consistent post-training pipeline to compare pretrained checkpoints.
@vkhoslaMost robotics companies claim internet video pre-training makes robots better. @RhodaAI actually tested it, on a real industrial task, scored the way a customer judges it, not a lab benchmark. Proud to back a team that proves it instead of asserting it. Great research article by Tongzhou Mu.
@RewkangAI Robotics research is showing that you can build increasingly powerful robot models by pretraining on the large preexisting corpus of web video data without collecting vast amounts of new data Visual dynamics understanding increases with more compute and that directly translates into more performant robot foundation models
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI

    16 Sources

    @tongzhou_mu1/ Setup. The video model is pre-trained on general web video, which contains no actions, then post-trained on robot demos of one industrial task, bearing unpacking. Each policy gets 100 to a few hundred robot trials, over 200 robot hours in total, every trial set up to a written procedure and audited afterwards. Evaluating on a real robot is hard, so the post describes how we do it in detail.
    @aliteracytongzhou led this entire work from start to finish (including training every model, finding bugs, and supervising every evaluation trial himself personally) - an absolute 🐐!! what excited me most about these results is that we can now use DINO FD as a predictor for post-training performance all the way from pre-training
    @startupjag@tongzhou_mu Scaling the AI model is one thing. Understanding whether it actually translates to better real-world robot performance is another. Fantastic work from @Tongzhou Mu and @RhodaAI putting it to the test.
    @RhodaAIToday marks week one of our new blog series. First one's from @tongzhou_mu . Will scaling web-video pre-training lead to better robot policies?
    @adam_patniRigorous evaluation on real world tasks is what will enable the best companies to outperform the rest. Scaling model size, compute, and improving web-data generation quality have empirically produced better robot policies across hundreds/thousands of trials at Rhoda. Major props to Tongzhou, who literally froze every single non-model thing in our office to produce this. I was literally maintaining months old infra to keep this eval going.
    @felixwyw@tongzhou_mu Congrats @tongzhou_mu amazing work!
    @GordonWetzstein@tongzhou_mu It's great to see a principled study on how pretraining at scale helps real-world robotics manipulation tasks. The insight is intuitive: pre-training on large web-video datasets helps downstream robotics tasks a lot!
    @charles_rqiGreat to see a principled study of video model pretraining for robot manipulation. I especially like the evaluation setup: a real-world industrial task and a strong, consistent post-training pipeline to compare pretrained checkpoints.
    @vkhoslaMost robotics companies claim internet video pre-training makes robots better. @RhodaAI actually tested it, on a real industrial task, scored the way a customer judges it, not a lab benchmark. Proud to back a team that proves it instead of asserting it. Great research article by Tongzhou Mu.
    @RewkangAI Robotics research is showing that you can build increasingly powerful robot models by pretraining on the large preexisting corpus of web video data without collecting vast amounts of new data Visual dynamics understanding increases with more compute and that directly translates into more performant robot foundation models
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet