ModAR introduces sequential prediction for robotics beyond RGB pixels
Its creators say ModAR predicts one modality at a time, with each prediction informing the next, to combine features capturing semantics, motion and geometry.
TLDR
ModAR’s creators argue that RGB pixels consume model capacity on fine details often irrelevant to robot policies. Their approach combines DINO features, point tracks and depth, predicting these different representations of the future in sequence. They report that their scratch-trained 30.1M model outperformed a 6B model initialized from a video model and fine-tuned on the same data.
Combined views
22
1 Source, first seen 4h ago
ModAR introduces sequential prediction for robotics beyond RGB pixels
Its creators say ModAR predicts one modality at a time, with each prediction informing the next, to combine features capturing semantics, motion and geometry.
TLDR
ModAR’s creators argue that RGB pixels consume model capacity on fine details often irrelevant to robot policies. Their approach combines DINO features, point tracks and depth, predicting these different representations of the future in sequence. They report that their scratch-trained 30.1M model outperformed a 6B model initialized from a video model and fine-tuned on the same data.