Users Compare Opus 5 and Sol on TB 3.0
X users compare Claude Opus 5 and Sol on Terminal-Bench 3.0 performance.
@teortaxesTex posted that TB 3.0 appears to reward both high intelligence and strong agentic behavior, adding that Opus 5 does not seem clearly better than Sol. @xlr8harder replied that Opus 5 can surface useful insights on problems where Sol stalls, yet it proves harder to steer and less consistent in follow-through. The reply noted that switching between the two models depending on the task might change which one appears stronger overall. Both comments treat direct head-to-head ranking as difficult to settle from available runs.
I'm confused by TB 3.0 looks like it measures general intelligence X general "agenticness", so both very strong and very harnessmaxxed models get ahead. I don't think Opus 5 is really superior to Sol
Terminal-Bench 3.0 最新榜单换王。这个评测让 AI Agent 在真实终端环境里完成任务,最后直接检查结果是否正确。 当前前三名分别是: 1. Claude Opus 5 Max + mini-SWE-agent:43.5% 2. GPT-5.6 Sol Max + Codex:34.6% 3. Claude Fable 5 + Claude Code:34.1% Terminal-Bench 3.0 在 7 月 23 日上线,目的就是把题重新做难。旧版前沿模型的成绩已经挤在一起,新版本希望重新拉开差距。首版有 74…

