Announcement
DiVeR preprint proposes prioritizing decision-critical states in robot verifier training
A researcher sharing the preprint says DiVeR gives greater training weight to moments when a robot’s candidate actions diverge.
TLDR
The researcher says DiVeR learns to rank a robot’s candidate actions from trajectory-level success or failure labels, without step-level annotations or extra environment interaction. They report average success rising from 58.3% to 71.9% across four real-robot tasks compared with single-sample inference. On RoboCasa, they also report over 700× faster verifier inference than a VLM-based baseline.
Combined views
827
2 Sources, first seen ago