Video World Model Trained on 15 Hours of Robot Video
Researchers trained the model on single-arm robot data, achieving cross-embodiment generalization via masked modeling techniques.
A video world model was trained using only 15 hours of footage from a single-arm robot. The model achieves zero-shot generalization to unseen embodiments. Related work on masked world modeling uses mask images as a unified interface to specify actions and tasks across different robot embodiments. The findings were shared and discussed by AI researchers including Jon Barron of Google DeepMind and Yilun Du of Harvard.
Combined views
1.5K
2 posts, first seen 3h ago