Frontier AI Models Score Nearly Identically on Coding Benchmark
Experts discuss non-monotonic performance curves and benchmark quirks for top models.
Entities: FrontierCode, Opus 5

A leaderboard screenshot for FrontierCode v1.1 shows two leading agentic coding systems at 53.4% and 53.5%. Commenters including Thomas Wolf, Jared Palmer, and Christopher Potts note that raising reasoning effort on Opus 5 reduces scores, unlike other models, and highlight effects from pure RL versus synthetic scaling. The discussion centers on why additional test-time compute sometimes hurts results and what this reveals about current evaluation limits.
This closely aligns with findings we report for Opus 4.6 in this post: https://bigspin.ai/resources/the-decline-of-token-level-purchasing-power
@Xinyu2ML the full graph makes this even weirder