Announcement
Claude Opus 5.5 leads the Hermes Bench index with a 63.31 mean score
NousResearch says the index covers four suites, with every model running the same harness and high reasoning where offered.
TLDR
NousResearch reports that Claude Opus 5.5 leads the Hermes Bench index with a 63.31 mean score and $4.99 mean cost per task. GPT 6 Astra follows at 56.25 and $11.61, then Sonnet 5.5 at 53.14 and $2.82. NousResearch says the index averages scores and costs per task across four suites.
Combined views
3.6K
1 Source, first seen ago
