• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    DIDO reportedly cuts robot action prediction time from 562 to 384 milliseconds

    A post about work by BAAI and collaborators describes DIDO’s approach: teach a one-step video model to retain the gripper and object details that can get lost when denoising is cut short.

    1 Source, 15d ago, first seen 15d ago

    TLDR

    A post describing work by BAAI and collaborators explains that World Action Models generate video predictions before deciding how a robot should act. It says backgrounds sharpen early during denoising, while the gripper, object and their interaction remain blurry until later steps. DIDO uses a four-step model to teach a single-step model to preserve those details. A pruning step then focuses computation on changing parts of the scene and compresses parts that stay static. The post reports action prediction in 384 milliseconds instead of 562.

    Combined views

    —

    1 Source, first seen 15d ago

    Combined views

    —

    1 Source, first seen 15d ago

    — likes
    — likes
    — comments
    — saves
    — reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    — comments
    — saves
    — reposts
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    1 Source

    @stepjamUKA robot's camera predicts the future to decide what to do next. The problem is that prediction takes time, and time is the one thing a robot moving in the real world doesn't have to spare. That's the tension a team from @BAAIBeijing and collaborators went after with World Action Models, systems that use video generation to imagine what happens next before deciding how to act. The video model builds that prediction through several rounds of denoising, and each round adds latency to the control loop. The obvious fix is to cut it down to one round. The team found out why that quietly breaks things. Watching the denoising process closely, they noticed the background of a scene sharpens almost immediately. The gripper, the object it's about to touch, and how the two interact stay blurry until several steps later. Cut the process short and you keep a crisp picture of the room and lose exactly the part the robot needed to act on. Their fix, called DIDO, trains a single denoising step to hold onto that missing detail. A slower four step model acts as a teacher, and the one step model is explicitly guided to track the object, the gripper, and their interaction directly, with extra supervision pointed at exactly those regions rather than the whole frame. A pruning step then spends compute where the scene is actually changing and compresses the parts that aren't. The result is a model that predicts an action in 384 milliseconds instead of 562, while doing just as well or better on manipulation benchmarks; 99.0 percent on LIBERO, and the strongest score among methods with no embodied pretraining on RoboTwin. It also held up on a real Galbot G1 arm across four tasks, including with clutter and repositioned objects added in. It's a good reminder that not every shortcut is free. Cutting a process down to save time can end up cutting away the part that made it work in the first place, and the only way to know is to look closely enough to see what's actually happening at each step. Paper: https://arxiv.org/abs/2609.15570 Project page: https://loveju1y.github.io/DIDO/