ECCV presenter says pretrained vision transformers enable fast real-world navigation
The presenter describes an approach that uses one number per image patch and pretrains a navigation policy with lidar input before removing lidar.
TLDR
The presenter of ECCV Poster #305 says a single numerical value per image patch from pretrained vision transformers enables fast real-world navigation. The approach, as described, includes a visual encoder distilled from different teacher models and a navigation policy pretrained with lidar input before lidar is removed.
Combined views
2.2K
2 Sources, first seen 19d ago