GPT-6-Astra outperforms prior OpenAI models on BullshitBench but trails Anthropic, a user says
The user estimates regrading all responses with newer models would cost about $3,000, but says changing judges only going forward would hurt continuity.
TLDR
A user reports that GPT-6-Astra did quite a bit better on BullshitBench than any previous OpenAI model, though it still fell short of Anthropic’s models. They say the judges are Sonnet 4.6, GPT-5.2 and Gemini 3.1 Pro, which they consider old. Regrading all responses with newer models would cost about $3,000, they estimate, while switching judges only going forward would hurt continuity. They suspect better judging could shift some of Astra’s amber grades to green or red.
Combined views
20.6K
1 Source, first seen 20d ago