Critics Flag Errors in Claude Opus 5 Chart
AI researchers highlight errors and misleading scales in Claude's performance table.
Anthropic's official Claude account posted a benchmark table claiming Opus 5 leads coding and knowledge evaluations. Multiple AI engineers and researchers replied with examples of chart errors, including selective highlighting of top scores and y-axis ranges set just above results. Miles Brundage framed the issue as over-reliance on Claude outputs rather than deliberate hype, while others noted repeated visual inaccuracies. The exchange underscores growing scrutiny of how frontier labs present model performance data. No corrections from Anthropic appear in the packet.
Combined views
11.4K
12 posts, first seen 2h ago