microSLAM team says it turns single-camera footage into 3D robot-training environments
The team says it built the system with ETH Zurich’s Computer Vision and Geometry Lab and that it ranks first on the LaMaria benchmark for monocular SLAM.
TLDR
The microSLAM team describes a system that maps surroundings and tracks camera position using a single RGB video stream. It says the resulting interactive 3D environments can be used to train behavioral foundation models. The team’s pitch is to capture robot-learning data through head-mounted footage of people doing everyday work, rather than fitting every worker with a camera rig. Its example: a person walks a factory floor, and a robot learns to navigate it.
Combined views
62K
3 Sources, first seen 20d ago