microSLAM team says it turns single-camera footage into 3D environments for robot training
The team says its system, built with ETH Zurich’s Computer Vision and Geometry Lab, ranks first on the LaMaria benchmark for monocular SLAM—mapping surroundings and tracking position with one camera.
TLDR
The team behind microSLAM says its system uses a single RGB video stream to create interactive 3D environments for training behavioral foundation models. Its announcement describes a use case where footage from someone walking a factory floor helps a robot learn to navigate it. The team argues that worker-worn cameras could make this data cheap to collect, with glasses scaling where camera rigs cannot. It also claims microSLAM holds up against systems using many more cameras and inertial measurement units.
Combined views
14
1 Source, first seen 20d ago