• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Reaction

    Reality Check launches a robot manipulation benchmark and leaderboard

    Its creators list 14,400 real-world rollouts across four models, with multiple tasks and data regimes.

    Lucas Beyer (bl16)LB
    Kosta DerpanisKD
    2 Sources, ,

    TLDR

    Reality Check’s creators announced the robot manipulation benchmark and leaderboard on September 29, listing 14,400 real-world rollouts across four models. A separate user praised the effort and said half the tasks are fully open, while half are held out to track potential benchmark gaming by future model versions.

    Combined views

    2.5K

    2 Sources, first seen 1h ago

    Combined views

    2.5K

    2 Sources, first seen 1h ago

    16 likes
    Today's Rank

    #16

    Today's Rank

    #16

    1h ago
    first seen 1h ago
    16 likes
    4 comments
    12 saves
    67 reposts
    4 comments
    12 saves
    67 reposts

    2 Sources

    Lucas Beyer (bl16)@giffmanaMissed this post the first time around, but i think this is a very cool and needed effort to thoroughly benchmark VLA and co. They build a leaderboard and half the tasks are fully open, half are held out to track potential benchmaxxing of future model versions.1h
    Kosta Derpanis@CSProfKGDRT @Nicolas_Keller: The Era of Evals is coming to AI robotics. Today we’re launching Reality Check: our leaderboard & first public robot m…30m
    Featured Source

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    2 Sources

    Lucas Beyer (bl16)@giffmanaMissed this post the first time around, but i think this is a very cool and needed effort to thoroughly benchmark VLA and co. They build a leaderboard and half the tasks are fully open, half are held out to track potential benchmaxxing of future model versions.1h
    Kosta Derpanis@CSProfKGDRT @Nicolas_Keller: The Era of Evals is coming to AI robotics. Today we’re launching Reality Check: our leaderboard & first public robot m…30m