Ox-Alpha Model Matches Claude Opus 4.8 on DeepSWE
Commenters compare Ox-Alpha results to GLM 5.3 and Kimi K3 on the same benchmark.
Zephyr posted surprise at the size of a model linked to a claim that Chinese Ox-Alpha reached 58.4 percent on the 113-task DeepSWE benchmark, close to Claude Opus 4.8 at 59 percent. TeortaxesTex replied that GLM 5.3 sits nearer Opus 5 on the eval and that Kimi K3 matched the Ox-Alpha score when released five weeks earlier. The reply noted Ox-Alpha stands out for separate reasons and urged Peter to adjust trend expectations.
Well, u are going to be extremely surprised by the size of the model
Them, careful analysis: this Chinese model might match Opus 4.8 Me, just looking at what the trend lines predict: China should have an Opus 4.8 model about now https://twitter.com/henryzhangumich/status/2091066210721141009