Opus 5.5 reportedly ranks second to seventh across GBENCH coding categories
A GBENCH tester says GPT-6-Astra remains comfortably ahead, while GPT-6-Sol performs within the margin of error of Opus 5.5. The tester cautions that the effort settings likely work against Opus 5.5.
TLDR
A GBENCH tester says they tested Opus 5.5, GPT-6-Sol and GPT-6-Luna across 100 open-ended coding and engineering environments with verifiable results. In their results, GPT-6-Astra remains comfortably ahead; Opus 5.5 ranks second to seventh across coding categories and second at producing code in one shot. GPT-6-Sol falls within the margin of error of Opus 5.5.
The tester says Opus 5.5 underperformed in their custom harness, frequently stopping when it judged its work good enough. They use adaptive/auto-reasoning where supported and “high” effort where configurable, a setup they say likely works against Opus 5.5.
Cost and speed comparisons also carry a caveat: the tester has used OpenAI Flex since Astra's release, describing it as half-price and a little slower. They caution that this matters when comparing with earlier OpenAI models on their charts.
