Ox-Alpha Model Matches Claude Opus 4.8 on DeepSWE
Commenters compare Ox-Alpha results to GLM 5.3 and Kimi K3 on the same benchmark.
TLDR
Zephyr posted surprise at the size of a model linked to a claim that Chinese Ox-Alpha reached 58.4 percent on the 113-task DeepSWE benchmark, close to Claude Opus 4.8 at 59 percent. TeortaxesTex replied that GLM 5.3 sits nearer Opus 5 on the eval and that Kimi K3 matched the Ox-Alpha score when released five weeks earlier. The reply noted Ox-Alpha stands out for separate reasons and urged Peter to adjust trend expectations.
Combined views
65.7K
3 Sources, first seen 39d ago