• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Mercor Shares GPT-6 Astra APEX-Accounting Scores

    Mercor posted benchmark results claiming GPT-6 Astra leads on APEX-Accounting tasks.

    ME
    1 Source, 25d ago, first seen 25d ago

    TLDR

    Mercor, an AI startup, posted that GPT-6 Astra achieved the top Pass@1 rate of 13.1 percent on the APEX-Accounting benchmark. The company stated Astra also reached a 60.0 percent mean score and outperformed GPT-5.6 Sol by 56 percent and Fable 5.1 by 12 percent in tasks passed. The post included a bar chart titled OpenAI GPT-6 As and explained that Pass@1 measures tasks scored fully correct at least once in four attempts. The tweet came from Mercor's official account.

    Combined views

    10.2K

    1 Source, first seen 25d ago

    Combined views

    10.2K

    1 Source, first seen 25d ago

    208 likes
    208 likes
    1 comments
    40 saves
    15 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    1 comments
    40 saves
    15 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @mercorGPT-6 Astra passes more tasks on APEX-Accounting than any other model. 13.1% Pass@1 (#1) 60.0% mean score (#2) Pass@1 is the proportion of tasks that a model scores 100% at least once across four attempts. Astra passes 56% more tasks than GPT-5.6 Sol and 12% more tasks than Fable 5.1. Most academic benchmarks measure model capabilities that are misaligned with real work. APEX benchmarks measure what enterprises actually care about. We built APEX-Accounting with @RampLabs to see if agents can handle a real company's books. Agents work the ledger in QuickBooks, tie it to bank statements in PDFs, chase figures across spreadsheets, and judge what is a real discrepancy. Results are graded against 2,186 criteria created by real accountants. See full leaderboard: https://www.mercor.com/apex/apex-accounting-leaderboard/

    1 Source

    @mercorGPT-6 Astra passes more tasks on APEX-Accounting than any other model. 13.1% Pass@1 (#1) 60.0% mean score (#2) Pass@1 is the proportion of tasks that a model scores 100% at least once across four attempts. Astra passes 56% more tasks than GPT-5.6 Sol and 12% more tasks than Fable 5.1. Most academic benchmarks measure model capabilities that are misaligned with real work. APEX benchmarks measure what enterprises actually care about. We built APEX-Accounting with @RampLabs to see if agents can handle a real company's books. Agents work the ledger in QuickBooks, tie it to bank statements in PDFs, chase figures across spreadsheets, and judge what is a real discrepancy. Results are graded against 2,186 criteria created by real accountants. See full leaderboard: https://www.mercor.com/apex/apex-accounting-leaderboard/