• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Rohan Paul Shares Cognition GPT-6 Astra FrontierCode Claim

    Rohan Paul posted Cognition's claim on GPT-6 Astra coding results.

    SW
    JA
    CO
    6 Sources, 27d ago, first seen 27d ago

    TLDR

    Rohan Paul, a Bengaluru-based machine learning engineer, posted that Cognition says GPT-6 Astra reached near-Fable 5 coding quality within 0.4 points at 64 percent lower rollout cost on FrontierCode 1.1. The post notes FrontierCode evaluates whether an AI coding agent can produce a pull request a maintainer would merge rather than stopping at functional correctness. A scatter plot chart titled Frontier accompanies the post. The statement attributes the performance data to Cognition without independent confirmation in the visible post.

    Combined views

    802.2K

    6 Sources, first seen 27d ago

    Combined views

    802.2K

    6 Sources, first seen 27d ago

    1.8K likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1.8K likes
    114 comments
    238 saves
    153 reposts
    114 comments
    238 saves
    153 reposts

    Sentiment

    Positive75.6%24.4%Negative

    Summary

    Sentiment

    Positive75.6%24.4%Negative

    Many accounts welcomed GPT-6 Astra's near-Fable 5 coding quality at 64% lower cost on FrontierCode for Devin, while others criticized limited accessibility and questioned whether the savings would hold in real use.

    Based on 46 sentiment-bearing replies from 45 accounts across 3 conversations.

    Summary

    Many accounts welcomed GPT-6 Astra's near-Fable 5 coding quality at 64% lower cost on FrontierCode for Devin, while others criticized limited accessibility and questioned whether the savings would hold in real use.

    Based on 46 sentiment-bearing replies from 45 accounts across 3 conversations.

    6 Sources

    @cognitionGPT-6 Astra is coming to Devin. On FrontierCode 1.1, Astra performs within 0.4 points of Fable 5 at a 64% lower cost. It also sets a new SOTA on our internal testing benchmark, generating more comprehensive tests, clearer reports, and better video evidence.
    @jxnlco64% lower cost per task!
    @rohanpaul_aiCognition says GPT-6 Astra reached near-Fable 5 coding quality (within 0.4 points) at 64% lower rollout cost on FrontierCode 1.1 Now, FrontierCode benchmark is quite unusual because it asks whether an AI coding agent can produce a pull request a maintainer would actually merge, rather than stopping at functional correctness. Cognition built 150 tasks with maintainers from 36 open-source repositories, grading correctness, regression safety, tests, scope, style, and adherence to each codebase's conventions. The score is a weighted rubric aggregate rather than a solve rate, and any run that misses a blocking requirement receives zero. METR separately found that roughly half of earlier SWE-bench Verified patches that passed automated tests still would not have been merged by maintainers. i.e. FrontierCode really measures review-quality code.
    @sherwinwu@LechMazur FrontierCode 1.1 by @cognition – 64% cheaper (against Fable 5, but 5.1 is similarly priced on the chart)

    6 Sources

    @cognitionGPT-6 Astra is coming to Devin. On FrontierCode 1.1, Astra performs within 0.4 points of Fable 5 at a 64% lower cost. It also sets a new SOTA on our internal testing benchmark, generating more comprehensive tests, clearer reports, and better video evidence.
    @jxnlco64% lower cost per task!
    @rohanpaul_aiCognition says GPT-6 Astra reached near-Fable 5 coding quality (within 0.4 points) at 64% lower rollout cost on FrontierCode 1.1 Now, FrontierCode benchmark is quite unusual because it asks whether an AI coding agent can produce a pull request a maintainer would actually merge, rather than stopping at functional correctness. Cognition built 150 tasks with maintainers from 36 open-source repositories, grading correctness, regression safety, tests, scope, style, and adherence to each codebase's conventions. The score is a weighted rubric aggregate rather than a solve rate, and any run that misses a blocking requirement receives zero. METR separately found that roughly half of earlier SWE-bench Verified patches that passed automated tests still would not have been merged by maintainers. i.e. FrontierCode really measures review-quality code.
    @sherwinwu@LechMazur FrontierCode 1.1 by @cognition – 64% cheaper (against Fable 5, but 5.1 is similarly priced on the chart)