DIDO reportedly cuts robot action prediction time from 562 to 384 milliseconds
A post about work by BAAI and collaborators describes DIDO’s approach: teach a one-step video model to retain the gripper and object details that can get lost when denoising is cut short.
TLDR
A post describing work by BAAI and collaborators explains that World Action Models generate video predictions before deciding how a robot should act. It says backgrounds sharpen early during denoising, while the gripper, object and their interaction remain blurry until later steps. DIDO uses a four-step model to teach a single-step model to preserve those details. A pruning step then focuses computation on changing parts of the scene and compresses parts that stay static. The post reports action prediction in 384 milliseconds instead of 562.
Combined views
—
1 Source, first seen 15d ago