Claude Opus 5 passes 23.9% of coding-agent benchmark evaluations, DAIR.AI says
DAIR.AI describes a 53-task benchmark across four domains that asks coding agents to deliver working customer service agents—not just code patches. It reports an expert human reference score of 82.2%.
TLDR
DAIR.AI reports that Claude Opus 5 running under Claude Code passes 23.9% of evaluations, compared with 82.2% for an expert human reference. It describes tasks that give a developer agent business records, a client who can answer questions, a production API and an inherited codebase, with hard limits on serving cost and model choice. The resulting customer service agents are scored by deploying them against held-out simulated users. DAIR.AI identifies shallow searches of business records, little client communication and limited experimentation with architecture or serving spend as failures—agents ship the first design that runs.
Combined views
9.8K
1 Source, first seen 23d ago