• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Ox-Alpha Model Matches Claude Opus 4.8 on DeepSWE

    Commenters compare Ox-Alpha results to GLM 5.3 and Kimi K3 on the same benchmark.

    T(
    ZE
    FB
    3 Sources, 39d ago, first seen 39d ago

    TLDR

    Zephyr posted surprise at the size of a model linked to a claim that Chinese Ox-Alpha reached 58.4 percent on the 113-task DeepSWE benchmark, close to Claude Opus 4.8 at 59 percent. TeortaxesTex replied that GLM 5.3 sits nearer Opus 5 on the eval and that Kimi K3 matched the Ox-Alpha score when released five weeks earlier. The reply noted Ox-Alpha stands out for separate reasons and urged Peter to adjust trend expectations.

    Combined views

    65.7K

    3 Sources, first seen 39d ago

    Combined views

    65.7K

    3 Sources, first seen 39d ago

    435 likes
    435 likes
    28 comments
    61 saves
    6 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Featured Source
    28 comments
    61 saves
    6 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    3 Sources

    @zephyr_z9Well, u are going to be extremely surprised by the size of the model
    @teortaxesTexThe problem is that on this same eval, GLM 5.3 is much closer to Opus 5 than to 4.8. And Kimi K3 scores the same and was released 5 weeks ago. "This Chinese model" (ox alpha) is surprising for somewhat different reasons. Calibrate your trend lines Peter
    @xeophon@teortaxesTex his own link puts K3 on 4.8 level

    3 Sources

    @zephyr_z9Well, u are going to be extremely surprised by the size of the model
    @teortaxesTexThe problem is that on this same eval, GLM 5.3 is much closer to Opus 5 than to 4.8. And Kimi K3 scores the same and was released 5 weeks ago. "This Chinese model" (ox alpha) is surprising for somewhat different reasons. Calibrate your trend lines Peter
    @xeophon@teortaxesTex his own link puts K3 on 4.8 level