• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Commentator Says GPT-7 Will Automate Blue Collar Work

    @scaling01, operator of LisanBench benchmark, made the statement about future AI models.

    RO
    DE
    AM
    18 Sources, 26d ago, first seen 26d ago

    TLDR

    @scaling01, who runs the LisanBench LLM reasoning benchmark, posted on X that GPT-7 will automate all blue collar work. The post quoted a thread by @chooi_jeq showing GPT-6 Astra reaching 95% on a robot control task versus 40% for Fable 5.1, using 6.2x fewer tokens at 2.3x lower cost. A retweet came from OpenAI researcher Aidan McLaughlin. The statement reflects the commentator's expressed view on model scaling rather than any verified development or product announcement.

    Combined views

    2.1M

    18 Sources, first seen 26d ago

    Combined views

    2.1M

    18 Sources, first seen 26d ago

    10.4K likes
    10.4K likes
    278 comments
    2.5K saves
    1.9K reposts
    278 comments
    2.5K saves
    1.9K reposts

    Sentiment

    Positive76.4%23.6%Negative

    Summary

    Many accounts welcomed GPT-6 Astra's jump to 95% on robot control tasks as real progress with fewer tokens and lower cost, while some replies called the benchmark misleading or raised safety concerns.

    Based on 122 sentiment-bearing replies from 110 accounts across 7 conversations.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    Sentiment

    Positive76.4%23.6%Negative

    Summary

    Many accounts welcomed GPT-6 Astra's jump to 95% on robot control tasks as real progress with fewer tokens and lower cost, while some replies called the benchmark misleading or raised safety concerns.

    Based on 122 sentiment-bearing replies from 110 accounts across 7 conversations.

    18 Sources

    @chooi_jeqGPT-6 Astra scored 95% on a robot control task, up from Fable 5.1's 40%, with 6.2x fewer output tokens at 2.3x lower cost. 🧵
    @scaling01world model shmord model GPT-7 is going to automate all blue collar work
    @YuXiang_IRVLI might be late to realize this, but the next frontier in robotics may be how to truly leverage advanced AI models for manipulation. Using them just as agents to call perception and planning modules isn’t enough. We need new ways to connect these models with the physical world.
    @aidan_mclauRT @chooi_jeq: GPT-6 Astra scored 95% on a robot control task, up from Fable 5.1's 40%, with 6.2x fewer output tokens at 2.3x lower cost. 🧵…
    @DJiafeiCould GPT-6 just API call MolmoAct2 checkpoint and ask put block into bowl?
    @BorisMPowerRobotics control might just be solved as a side quest of improving LLMs - pretty exciting! And it’ll be a lot more general
    @DorialexanderMeanwhile we had an entire discourse cycle in Europe that LLMs were saturated and it was high time to invest in non-LLMs leapfrogs for robotics.
    @chris_j_paxtonA long way to go before these models can do control tasks, but genuinely real progress here
    @teortaxesTexRT @Dorialexander: Meanwhile we had an entire discourse cycle in Europe that LLMs were saturated and it was high time to invest in non-LLMs…
    @mmmbchang🤯

    18 Sources

    @chooi_jeqGPT-6 Astra scored 95% on a robot control task, up from Fable 5.1's 40%, with 6.2x fewer output tokens at 2.3x lower cost. 🧵
    @scaling01world model shmord model GPT-7 is going to automate all blue collar work
    @YuXiang_IRVLI might be late to realize this, but the next frontier in robotics may be how to truly leverage advanced AI models for manipulation. Using them just as agents to call perception and planning modules isn’t enough. We need new ways to connect these models with the physical world.
    @aidan_mclauRT @chooi_jeq: GPT-6 Astra scored 95% on a robot control task, up from Fable 5.1's 40%, with 6.2x fewer output tokens at 2.3x lower cost. 🧵…
    @DJiafeiCould GPT-6 just API call MolmoAct2 checkpoint and ask put block into bowl?
    @BorisMPowerRobotics control might just be solved as a side quest of improving LLMs - pretty exciting! And it’ll be a lot more general
    @DorialexanderMeanwhile we had an entire discourse cycle in Europe that LLMs were saturated and it was high time to invest in non-LLMs leapfrogs for robotics.
    @chris_j_paxtonA long way to go before these models can do control tasks, but genuinely real progress here
    @teortaxesTexRT @Dorialexander: Meanwhile we had an entire discourse cycle in Europe that LLMs were saturated and it was high time to invest in non-LLMs…
    @mmmbchang🤯