Grok 4.5 Outperforms GPT-5.5 and Claude Opus 4.8 in Workplace Benchmarks
Reactions from ranked influencers
4 postsTry Grok Build!
Grok 4.5 is leading on actual professional work In Snorkel AI’s benchmark, Grok 4.5 combined with Grok Build was tested against GPT 5.5 and Claude Opus 4.8 across nearly 2,000 expert-created workplace tasks involving real documents, spreadsheets, presentations and professional analysis Grok outperformed both GPT 5.5 and Claude Opus 4.8 overall, while leading by even wider margins across several high-judgment fields: • Education: 58% • Legal work: 40% • Quality assurance: 37% • Healthcare: 35% It also recorded the lowest failure rate across every error category Snorkel measured, including missing analysis, incorrect recommendations, poor structure and missing sources The most important result is not simply that Grok completed more tasks It produced better professional deliverables, made fewer critical mistakes and provided more specific, actionable recommendations where competing models often returned generic work Grok 4.5 is proving that real-world usefulness matters more than benchmark scores alone
Try Grok Build, it’s awesome! http://X.ai/cli
Grok 4.5 now also ranked #1 on the Long-Horizon Terminal-Bench by binary pass rate, outperforming Claude Fable 5, Claude Opus 4.8 and GPT-5.6-sol Under the strictest scoring metric - where a task counts only if it is fully solved with a perfect reward and zero errors......Grok 4.5 finished clearly ahead of every other tested model This matters for real-world coding, automation and difficult engineering work because long-horizon terminal tasks require much more than making partial progress The model has to maintain context, recover from mistakes and successfully complete an entire workflow across hundreds of steps Grok 4.5 is showing serious strength on complex agentic tasks over extended periods of time
Grok is a solid workhorse
Still haven’t found a real use case where Grok 4.5 fails at something I ask vs GPT-5.6 Sol, Fable 5, or Kimi K3. All of them are excellent — but Grok 4.5 wins on price and speed by far.
I’m currently using Grok 4.5 99% of the time now for both planning and tasks. The fact that I don’t have to switch out to a cheaper model to do tasks means there is less context window model switching and faster speed. Feels incredible.
We had a team of agents rebuild SQLite from its 835-page manual. It created a replica in Rust which passed 100% of a held-out test suite. Interestingly, cost varied 15x depending on which model mix we used.
Combined views
15.1M
4 posts, first seen 1d ago