• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Robot data estimate is off by 100–500 times, a user argues

    The user says the cited figures—64TB of robot data per day versus a 120TB language-model training corpus—imply roughly half as much data, not 200 times as much.

    ZE
    1 Source, 29d ago, first seen 29d ago

    TLDR

    A user disputes a headline they describe as claiming that a robot may generate 200 times the data in an entire frontier language-model training corpus in one day. They say the cited figures are 64TB per day for the robot and 120TB for the corpus, contradicting that comparison. The user also challenges the underlying data estimates, saying the cited research uses the DROID database, whose 350 hours of training data total 1.7TB, with the raw footage totaling 8.7TB.

    Combined views

    33.3K

    1 Source, first seen 29d ago

    Combined views

    33.3K

    1 Source, first seen 29d ago

    115 likes
    115 likes
    4 comments
    32 saves
    3 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    4 comments
    32 saves
    3 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @zephyr_z9The first error from the sell-side guy The headline says a robot may generate 200x more data in "1 day" vs the entire frontier LLM corpus If u look at his own data, the LLM training corpus is 120TB, while the robot generates 64TB per day The robot is generating 0.5x data per day compared to the LLM training corpus (although training corpus these days are much bigger than 120TB) Secondly, the research he is citing uses the DROID database The actual size of the 350 hours of training data 1.7TB and the raw 350-hour footage is 8.7TB So, the sell-side bro is off by 100x-500x

    1 Source

    @zephyr_z9The first error from the sell-side guy The headline says a robot may generate 200x more data in "1 day" vs the entire frontier LLM corpus If u look at his own data, the LLM training corpus is 120TB, while the robot generates 64TB per day The robot is generating 0.5x data per day compared to the LLM training corpus (although training corpus these days are much bigger than 120TB) Secondly, the research he is citing uses the DROID database The actual size of the 350 hours of training data 1.7TB and the raw 350-hour footage is 8.7TB So, the sell-side bro is off by 100x-500x