• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Scoble Retweets Claim GPT-6 Astra Beats Fable 5.1

    Replies post benchmark tables and charts comparing the models on internal tasks.

    RA
    PW
    BP
    49 Sources, 27d ago, first seen 27d ago

    TLDR

    Robert Scoble retweeted ChrisGPT stating that GPT-6 Astra outperformed Claude Fable 5.1 on Terminal science and Automation Bench. Separate posts by Andrew Curran, scaling01, nrehiew_, and kimmonismus include screenshots of tables and line charts for Data Science Tasks, DeepSWE v1.1, FrontierCode 1.1 Extended, Design Tasks, and ARGI-AGI 3. The images list GPT-6 Astra alongside GPT-5.6 Sol[2], Claude Fable 5.1, Claude Fable 5, Claude Opus 5, and Gemini 3.8 Flash. kimmonismus also wrote that OpenAI crushed Anthropic’s IPO with the results. All content consists of social media claims and attached media.

    Combined views

    4.7M

    49 Sources, first seen 27d ago

    Combined views

    4.7M

    49 Sources, first seen 27d ago

    29.4K likes
    29.4K likes
    1.3K comments
    4.3K saves
    1.3K reposts
    1.3K comments
    4.3K saves
    1.3K reposts

    Sentiment

    Positive41.3%58.7%Negative

    Summary

    Sentiment

    Positive41.3%58.7%Negative

    Positive accounts praised GPT-6 Astra's high benchmark scores such as 98.6% on ARC-AGI-3 as groundbreaking, while negative replies dismissed the results over non-standard harnesses and called the benchmarks meaningless.

    Based on 222 sentiment-bearing replies from 201 accounts across 13 conversations.

    Summary

    Positive accounts praised GPT-6 Astra's high benchmark scores such as 98.6% on ARC-AGI-3 as groundbreaking, while negative replies dismissed the results over non-standard harnesses and called the benchmarks meaningless.

    Based on 222 sentiment-bearing replies from 201 accounts across 13 conversations.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    49 Sources

    @scaling01are you ready for a nuke to hit ARC-AGI-3?
    @kimmonismusGPT-6 Astra Benchmarks: FrontierMath Tier 4 v2: 97.6% DeepSWE v1.1: 74.1% BenchCAD: 95.9% GPQA Diamond: 96% ExploitBench: 100% ARC-AGI-3: 98.6% (reported) Access in "plans to bring the new model to paying ChatGPT subscribers and its API over the next several days."
    @ChrisGPTOh my God it MOGGED Fable 5.1 on Terminal science and on Automation Bench😭
    @ScobleizerRT @ChrisGPT: Oh my God it MOGGED Fable 5.1 on Terminal science and on Automation Bench😭
    @Yampeleg98.6% on ARC-AGI-3 is nuts
    @teortaxesTexwhoever dunked on this: I generously advise you again to trust my gut more I don't give a fuck about your reasoned arguments "it has been revealed to me in a dream" and "I feel like this Indian anon is legit" work well enough for me
    @AndrewCurran_Benchmarks and other information getting leaked by the media, this is from Venturebeat. Possibly the original press embargo has passed due to the delay. I think it was supposed to go live at 11am.
    @reiinakanoexploitbench 100% so is this with or without hacking into huggingface's servers?
    @nrehiew_Interesting that of all benches only the coding benchmarks DeepSWE and FrontierCode report using number of output tokens instead of API cost
    @_arohan_AGI is 74% deepswe.

    49 Sources

    @scaling01are you ready for a nuke to hit ARC-AGI-3?
    @kimmonismusGPT-6 Astra Benchmarks: FrontierMath Tier 4 v2: 97.6% DeepSWE v1.1: 74.1% BenchCAD: 95.9% GPQA Diamond: 96% ExploitBench: 100% ARC-AGI-3: 98.6% (reported) Access in "plans to bring the new model to paying ChatGPT subscribers and its API over the next several days."
    @ChrisGPTOh my God it MOGGED Fable 5.1 on Terminal science and on Automation Bench😭
    @ScobleizerRT @ChrisGPT: Oh my God it MOGGED Fable 5.1 on Terminal science and on Automation Bench😭
    @Yampeleg98.6% on ARC-AGI-3 is nuts
    @teortaxesTexwhoever dunked on this: I generously advise you again to trust my gut more I don't give a fuck about your reasoned arguments "it has been revealed to me in a dream" and "I feel like this Indian anon is legit" work well enough for me
    @AndrewCurran_Benchmarks and other information getting leaked by the media, this is from Venturebeat. Possibly the original press embargo has passed due to the delay. I think it was supposed to go live at 11am.
    @reiinakanoexploitbench 100% so is this with or without hacking into huggingface's servers?
    @nrehiew_Interesting that of all benches only the coding benchmarks DeepSWE and FrontierCode report using number of output tokens instead of API cost
    @_arohan_AGI is 74% deepswe.