Frontier AI Models Score Nearly Identically on Coding Benchmark
Experts note quirks in FrontierCode benchmark affecting scores for models like Opus 5.
Entities: FrontierCode, Opus 5

Discussions among AI researchers and engineers reveal non-monotonic performance on the FrontierCode agentic coding benchmark. Several frontier models, including Anthropic's Opus 5, show scores declining as reasoning effort increases beyond medium levels. The benchmark evaluates mergeability using maintainer heuristics such as scope limits rather than pure task completion, creating noisy results and penalizing unprompted refactors. Similar patterns appear across multiple models, contrasting with steadier curves from Fable and Sol. Observations align with prior reports on token-level purchasing power and highlight challenges in industrializing test-time scaling approaches.
This closely aligns with findings we report for Opus 4.6 in this post: https://bigspin.ai/resources/the-decline-of-token-level-purchasing-power
@Xinyu2ML the full graph makes this even weirder