GPT-6-Astra beats earlier OpenAI models but trails Anthropic on BullshitBench, a user reports
The user estimates that regrading all responses with newer AI judges would cost about $3,000, but says switching judges only going forward would hurt continuity.
TLDR
A user reports that GPT-6-Astra did quite a bit better than previous OpenAI models on BullshitBench, but did not quite reach Anthropic models. They say the benchmark’s judges—Sonnet 4.6, GPT-5.2 and Gemini 3.1 Pro—are pretty old. Ideally, they would regrade all responses with Fable- or Astra-level models, at an estimated cost of about $3,000. Switching judges only going forward would hurt continuity, they say, while better judging might change some of Astra’s amber grades to green or red.
Combined views
209
1 Source, first seen 20d ago