• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    microSLAM team says it turns single-camera footage into 3D robot-training environments

    The team says it built the system with ETH Zurich’s Computer Vision and Geometry Lab and that it ranks first on the LaMaria benchmark for monocular SLAM.

    CP
    GH
    VN
    3 Sources, ,

    TLDR

    The microSLAM team describes a system that maps surroundings and tracks camera position using a single RGB video stream. It says the resulting interactive 3D environments can be used to train behavioral foundation models. The team’s pitch is to capture robot-learning data through head-mounted footage of people doing everyday work, rather than fitting every worker with a camera rig. Its example: a person walks a factory floor, and a robot learns to navigate it.

    Combined views

    62K

    3 Sources, first seen 20d ago

    Combined views

    62K

    3 Sources, first seen 20d ago

    596 likes
    20d ago
    first seen 20d ago
    596 likes
    22 comments
    454 saves
    55 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    22 comments
    454 saves
    55 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    3 Sources

    @notgianneiToday we're announcing microSLAM, a monocular SLAM system built by our computer vision team in Zürich with ETH Zurich's Computer Vision and Geometry Lab. It ranks first on the LaMaria benchmark for monocular SLAM, and holds up against systems that carry many more cameras and IMUs. Ours runs on a single RGB stream. Robot learning is bottlenecked on data, and the largest untapped source is people going about their work. A camera on someone's head for an afternoon in a plant is a record of how that plant actually operates, including all the parts nobody writes down. That data is only useful if you can recover the geometry, and geometry from one moving camera in a world that will not sit still is the hard version of the problem. Solving it monocular is what makes the data cheap. Rigs do not scale to every worker. Glasses do. microSLAM turns that footage into interactive 3D environments where behavioral foundation models can be trained. One person walks a factory floor. A robot learns to navigate it.
    @_varunnairCongratulations @microagi on microSLAM I’ve worked with plain monocular RGB to get pose estimation and it is a hard problem to get the last 10% right. Excited to see the Physical AI Data evolve and earn the right to solve the next problem.
    @chris_j_paxtonRT @_varunnair: Congratulations @microagi on microSLAM I’ve worked with plain monocular RGB to get pose estimation and it is a hard proble…

    3 Sources

    @notgianneiToday we're announcing microSLAM, a monocular SLAM system built by our computer vision team in Zürich with ETH Zurich's Computer Vision and Geometry Lab. It ranks first on the LaMaria benchmark for monocular SLAM, and holds up against systems that carry many more cameras and IMUs. Ours runs on a single RGB stream. Robot learning is bottlenecked on data, and the largest untapped source is people going about their work. A camera on someone's head for an afternoon in a plant is a record of how that plant actually operates, including all the parts nobody writes down. That data is only useful if you can recover the geometry, and geometry from one moving camera in a world that will not sit still is the hard version of the problem. Solving it monocular is what makes the data cheap. Rigs do not scale to every worker. Glasses do. microSLAM turns that footage into interactive 3D environments where behavioral foundation models can be trained. One person walks a factory floor. A robot learns to navigate it.
    @_varunnairCongratulations @microagi on microSLAM I’ve worked with plain monocular RGB to get pose estimation and it is a hard problem to get the last 10% right. Excited to see the Physical AI Data evolve and earn the right to solve the next problem.
    @chris_j_paxtonRT @_varunnair: Congratulations @microagi on microSLAM I’ve worked with plain monocular RGB to get pose estimation and it is a hard proble…