• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    GPT-6-Astra beats earlier OpenAI models but trails Anthropic on BullshitBench, a user reports

    The user estimates that regrading all responses with newer AI judges would cost about $3,000, but says switching judges only going forward would hurt continuity.

    LA
    1 Source, 20d ago, first seen 20d ago

    TLDR

    A user reports that GPT-6-Astra did quite a bit better than previous OpenAI models on BullshitBench, but did not quite reach Anthropic models. They say the benchmark’s judges—Sonnet 4.6, GPT-5.2 and Gemini 3.1 Pro—are pretty old. Ideally, they would regrade all responses with Fable- or Astra-level models, at an estimated cost of about $3,000. Switching judges only going forward would hurt continuity, they say, while better judging might change some of Astra’s amber grades to green or red.

    Combined views

    209

    1 Source, first seen 20d ago

    Combined views

    209

    1 Source, first seen 20d ago

    4 reposts
    4 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @scaling01RT @petergostev: BullshitBench: GPT-6-Astra did quite a bit better than any previous OpenAI models but not quite reaching Anthropic models…

    1 Source

    @scaling01RT @petergostev: BullshitBench: GPT-6-Astra did quite a bit better than any previous OpenAI models but not quite reaching Anthropic models…