• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    ModAR introduces sequential prediction for robotics beyond RGB pixels

    Its creators say ModAR predicts one modality at a time, with each prediction informing the next, to combine features capturing semantics, motion and geometry.

    Mariya I. VasilevaMI
    Adam HungAH
    2 Sources, ,

    TLDR

    ModAR’s creators argue that RGB pixels consume model capacity on fine details often irrelevant to robot policies. Their approach combines DINO features, point tracks and depth, predicting these different representations of the future in sequence. They report that their scratch-trained 30.1M model outperformed a 6B model initialized from a video model and fine-tuned on the same data.

    Combined views

    37.6K

    2 Sources, first seen 21d ago

    Combined views

    37.6K

    2 Sources, first seen 21d ago

    406 likes
    21d ago
    first seen 21d ago
    406 likes
    13 comments
    352 saves
    117 reposts
    13 comments
    352 saves
    117 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    Adam Hung@AdamjhungWorld-action models typically imagine the future in RGB – but are pixels really the right representation for robotics? Our bet is no: RGB spends capacity on fine-grained details and variation that are often irrelevant for robot policies. DINO features, point tracks, and depth capture more useful features like semantics, motion, and geometry. But no single modality captures everything — how can we effectively combine them? We introduce ✨ModAR✨, which predicts the future one modality at a time, with each prediction informing the next. We train from scratch and find that this formulation performs best. Our 30.1M scratch-trained model even outperforms a 6B video-model-initialized model finetuned on the same data! 🧵 [1/8]21d
    Mariya I. Vasileva@mariyaivasilevaRT @Adamjhung: World-action models typically imagine the future in RGB – but are pixels really the right representation for robotics? Our…20d

    2 Sources

    Adam Hung@AdamjhungWorld-action models typically imagine the future in RGB – but are pixels really the right representation for robotics? Our bet is no: RGB spends capacity on fine-grained details and variation that are often irrelevant for robot policies. DINO features, point tracks, and depth capture more useful features like semantics, motion, and geometry. But no single modality captures everything — how can we effectively combine them? We introduce ✨ModAR✨, which predicts the future one modality at a time, with each prediction informing the next. We train from scratch and find that this formulation performs best. Our 30.1M scratch-trained model even outperforms a 6B video-model-initialized model finetuned on the same data! 🧵 [1/8]21d
    Mariya I. Vasileva@mariyaivasilevaRT @Adamjhung: World-action models typically imagine the future in RGB – but are pixels really the right representation for robotics? Our…20d