• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Justin Johnson Predicts Robot Training From Five Photos

    MTS post quotes World Labs co-founder Justin Johnson on robot training via five phone photos.

    FL
    A1
    EL
    4 Sources, 29d ago, first seen 29d ago

    TLDR

    @MTSlive posted a quote from World Labs co-founder @jcjohnss describing a real-to-sim-to-real approach. Johnson outlined in-context learning where a robotics foundation model could adapt after one demonstration of a task. He suggested Atlas might train for any new environment in minutes using five photos taken on a phone. The post includes an 86-second video clip of Johnson seated at a table during the interview.

    Combined views

    85.1K

    4 Sources, first seen 29d ago

    Combined views

    85.1K

    4 Sources, first seen 29d ago

    432 likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    432 likes
    30 comments
    177 saves
    50 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    30 comments
    177 saves
    50 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    4 Sources

    @omarsar0Omni models are the next frontier. Simply put it, this is the most exciting release I've seen this year. This work is so ahead there isn't even a benchmark to measure the general capabilities of these type of world models. Scaling seems unlocked too. Wow!
    @MTSliveWorld Labs co-founder @jcjohnss predicts a real-to-sim-to-real future where Atlas could train a robot for any new environment in minutes using just 5 photos from your phone: "One is this notion of fully in-context learning. Maybe I've got a robotics foundation model, then I can demonstrate the robot once how a task should be performed, and that's enough for the robot to figure out." "Another version is real to sim to real. Maybe I've got my pre-trained robotics foundation model, but I want to adapt it to this particular environment, like this studio, this table, moving this microphone around." "There's a version where you could come in here and take just a couple casual videos with your phone or a couple images with your phone, use Atlas to reconstruct the space, and now stage all kinds of robotics interactions in this studio space." "If you could lower the barrier to doing that, you could do it in five minutes to take five photos, upload it to Atlas, have it generate a simulation, maybe describe in natural language what kind of task you want the robot to do, have an agent go build that simulation for you, RL fine-tune your general purpose robotics foundation model, and now you've got a robot maybe in the span of a couple minutes that could come in and is perfectly adapted to this space." @theworldlabs
    @drfeifeiRT @MTSlive: World Labs co-founder @jcjohnss predicts a real-to-sim-to-real future where Atlas could train a robot for any new environment…
    @a16zWorld Labs co-founder Justin Johnson on how five photos from your phone could teach a robot to work in a room it has never seen: "I think there's a couple different tech trees that people are working on. One is this notion of fully in-context learning. Maybe I've got a robotics foundation model, then I can demonstrate to the robot once how a task should be performed, and that's enough for the robot to figure out that task." "Another version is what we're calling real-to-sim-to-real. Maybe I've got my pre-trained robotics foundation model, but I want to adapt it to this particular environment." "Like this studio... You could come in here and take just a couple casual videos with your phone or a couple images with your phone, use Atlas to reconstruct the space, and now stage all kinds of robotics interactions in this studio space in particular." "Literally you could do it in five minutes. Take five photos, upload it to Atlas, have it generate a simulation, describe in natural language what kind of task you want the robot to do, have an agent go build that simulation for you, RL fine-tune your general purpose robotics foundation model, and now you've got a robot... that's perfectly adapted to this space." @jcjohnss on @MTSlive

    4 Sources

    @omarsar0Omni models are the next frontier. Simply put it, this is the most exciting release I've seen this year. This work is so ahead there isn't even a benchmark to measure the general capabilities of these type of world models. Scaling seems unlocked too. Wow!
    @MTSliveWorld Labs co-founder @jcjohnss predicts a real-to-sim-to-real future where Atlas could train a robot for any new environment in minutes using just 5 photos from your phone: "One is this notion of fully in-context learning. Maybe I've got a robotics foundation model, then I can demonstrate the robot once how a task should be performed, and that's enough for the robot to figure out." "Another version is real to sim to real. Maybe I've got my pre-trained robotics foundation model, but I want to adapt it to this particular environment, like this studio, this table, moving this microphone around." "There's a version where you could come in here and take just a couple casual videos with your phone or a couple images with your phone, use Atlas to reconstruct the space, and now stage all kinds of robotics interactions in this studio space." "If you could lower the barrier to doing that, you could do it in five minutes to take five photos, upload it to Atlas, have it generate a simulation, maybe describe in natural language what kind of task you want the robot to do, have an agent go build that simulation for you, RL fine-tune your general purpose robotics foundation model, and now you've got a robot maybe in the span of a couple minutes that could come in and is perfectly adapted to this space." @theworldlabs
    @drfeifeiRT @MTSlive: World Labs co-founder @jcjohnss predicts a real-to-sim-to-real future where Atlas could train a robot for any new environment…
    @a16zWorld Labs co-founder Justin Johnson on how five photos from your phone could teach a robot to work in a room it has never seen: "I think there's a couple different tech trees that people are working on. One is this notion of fully in-context learning. Maybe I've got a robotics foundation model, then I can demonstrate to the robot once how a task should be performed, and that's enough for the robot to figure out that task." "Another version is what we're calling real-to-sim-to-real. Maybe I've got my pre-trained robotics foundation model, but I want to adapt it to this particular environment." "Like this studio... You could come in here and take just a couple casual videos with your phone or a couple images with your phone, use Atlas to reconstruct the space, and now stage all kinds of robotics interactions in this studio space in particular." "Literally you could do it in five minutes. Take five photos, upload it to Atlas, have it generate a simulation, describe in natural language what kind of task you want the robot to do, have an agent go build that simulation for you, RL fine-tune your general purpose robotics foundation model, and now you've got a robot... that's perfectly adapted to this space." @jcjohnss on @MTSlive