• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Announcement

    DiVeR preprint proposes prioritizing decision-critical states in robot verifier training

    A researcher sharing the preprint says DiVeR gives greater training weight to moments when a robot’s candidate actions diverge.

    Sharon LiSL
    Seongheon ParkSP
    2 Sources, ,

    TLDR

    The researcher says DiVeR learns to rank a robot’s candidate actions from trajectory-level success or failure labels, without step-level annotations or extra environment interaction. They report average success rising from 58.3% to 71.9% across four real-robot tasks compared with single-sample inference. On RoboCasa, they also report over 700× faster verifier inference than a VLM-based baseline.

    Combined views

    827

    2 Sources, first seen 3h ago

    Combined views

    827

    2 Sources, first seen 3h ago

    15 likes
    3h ago
    first seen 3h ago
    15 likes
    1 comments
    3 saves
    16 reposts
    1 comments
    3 saves
    16 reposts
    Featured Source

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    2 Sources

    Seongheon Park@seongheon_96How can we learn better verifiers to make the Generate → Verify → Act loop more effective through Best-of-N action selection in Vision-Language-Action (VLA) models? Excited to share our new preprint, DiVeR: Decision-Critical Verifier Learning for VLA Test-Time Scaling 🎉 The idea starts with a simple observation: not all states matter equally for verifier learning. During routine motion, such as moving through free space or carrying an object, candidate actions tend to be similar. Choosing among them makes less difference, so these states provide limited signal for learning which actions are better. But at a few decision-critical moments, such as aligning a grasp or placement, candidates can diverge—and even a small difference in action choice can change the outcome. We built DiVeR around two ideas: • Identify decision-critical states by measuring disagreement among sampled action representations. • Focus verifier learning on those states by giving them greater training weight. At test time, Generate → Verify → Act repeats at every decision step: the frozen VLA samples multiple action chunks, DiVeR ranks them, and the robot executes the highest-scoring chunk. The process then repeats with the next observation. The verifier learns from trajectory-level success/failure labels, without requiring step-level annotations or additional environment interaction. Across four real-robot tasks, average success rises from 58.3% to 71.9% compared with single-sample inference. On RoboCasa, DiVeR also achieves over 700× faster verifier inference than the VLM-based baseline. This work was done during my wonderful time at MSRA-Tokyo. A big thank-you to Heecheol Kim, @shulin_tian , Lilika Makabe, @namikosaito_ , Katsushi Ikeuchi, Yasuyuki Matsushita, and my advisor @SharonYixuanLi for all the thoughtful discussions, guidance, and support! The work has also been accepted to the PTA Workshop at NeurIPS 2026! (https://ptaworkshop.github.io/) #Robotics #EmbodiedAI #VLA #NeurIPS3h
    Sharon Li@SharonYixuanLiRT @seongheon_96: How can we learn better verifiers to make the Generate → Verify → Act loop more effective through Best-of-N action select…1h
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    Seongheon Park@seongheon_96How can we learn better verifiers to make the Generate → Verify → Act loop more effective through Best-of-N action selection in Vision-Language-Action (VLA) models? Excited to share our new preprint, DiVeR: Decision-Critical Verifier Learning for VLA Test-Time Scaling 🎉 The idea starts with a simple observation: not all states matter equally for verifier learning. During routine motion, such as moving through free space or carrying an object, candidate actions tend to be similar. Choosing among them makes less difference, so these states provide limited signal for learning which actions are better. But at a few decision-critical moments, such as aligning a grasp or placement, candidates can diverge—and even a small difference in action choice can change the outcome. We built DiVeR around two ideas: • Identify decision-critical states by measuring disagreement among sampled action representations. • Focus verifier learning on those states by giving them greater training weight. At test time, Generate → Verify → Act repeats at every decision step: the frozen VLA samples multiple action chunks, DiVeR ranks them, and the robot executes the highest-scoring chunk. The process then repeats with the next observation. The verifier learns from trajectory-level success/failure labels, without requiring step-level annotations or additional environment interaction. Across four real-robot tasks, average success rises from 58.3% to 71.9% compared with single-sample inference. On RoboCasa, DiVeR also achieves over 700× faster verifier inference than the VLM-based baseline. This work was done during my wonderful time at MSRA-Tokyo. A big thank-you to Heecheol Kim, @shulin_tian , Lilika Makabe, @namikosaito_ , Katsushi Ikeuchi, Yasuyuki Matsushita, and my advisor @SharonYixuanLi for all the thoughtful discussions, guidance, and support! The work has also been accepted to the PTA Workshop at NeurIPS 2026! (https://ptaworkshop.github.io/) #Robotics #EmbodiedAI #VLA #NeurIPS3h
    Sharon Li@SharonYixuanLiRT @seongheon_96: How can we learn better verifiers to make the Generate → Verify → Act loop more effective through Best-of-N action select…1h