• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Ox Alpha Scores 63 Percent on DeepSWE Subset

    Posts discuss Ox Alpha results on DeepSWE benchmark and comparisons to other models.

    T(
    T-
    CH
    5 Sources, 40d ago, first seen 40d ago

    TLDR

    A retweet from @theo shares @winkey_h noting Ox Alpha reached around 63 percent at 47K average output tokens on a DeepSWE subset, calling the result pareto optimal. A quote from @kimmonismus states that @davis7 fully tested Ox Alpha on the DeepSWE set, with overall performance described as more or less on par with GPT-5.6 Sol mid. The same post speculates the model may be GLM-5.3 Flash capable of local runs on hardware such as a DGX Spark. One generated headline in the packet refers to Ox Alpha as a stealth model from Opencode with 1M context length and multimodal support.

    Combined views

    547.3K

    5 Sources, first seen 40d ago

    Combined views

    547.3K

    5 Sources, first seen 40d ago

    6.8K likes
    6.8K likes
    347 comments
    1K saves
    222 reposts
    347 comments
    1K saves
    222 reposts

    Sentiment

    Positive78.1%21.9%Negative

    Summary

    Many accounts welcomed Ox Alpha's near-GPT-5.6 Sol Mid performance on DeepSWE because it points to capable open models usable locally, while some questioned the evaluation subsets and production latency.

    Based on 33 sentiment-bearing replies from 32 accounts across 2 conversations.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    5 Sources

    @theoRT @winkey_h: Ok ox-alpha has that big model smell I'm getting around 63% at 47K avg output tokens on a DeepSWE subset This is pareto opt…
    @kimmonismusOx Alpha has now been fully tested against the DeepSWE set by @davis7 . Overall, it's performing more or less on par with GPT-5.6 Sol mid. But he makes a good point. If it really is GLM-5.3 Flash that's now performing at the level of 5.6 Sol mid and could run locally on a DGX Spark, I wouldn't just be satisfied, it would be an absolutely fantastic deal. 5.6 Sol mid running locally, with the only cost being power consumption, for all sorts of tasks would be incredibly great. This model 24/7 hermes local: game changer.
    @teortaxesTexalternative results I think 60-61% is realistic
    @thdxrlot of guesses on what ox alpha is but they are all wrong, kinda disappointed so just going to tell you ox alpha is a new kind of llm that recursively updates a persistent latent state instead of reasoning entirely through tokens this lets internal representations converge before anything is actually decoded those attractors generate shards that encode transformations between latent states rather than the states themselves at sufficient density these shards compose into metaparameters that dynamically alter the residual geometry of the model without changing its weights. we built it because there was one thing simply too large to fit inside the context window of any existing model your mom

    5 Sources

    @theoRT @winkey_h: Ok ox-alpha has that big model smell I'm getting around 63% at 47K avg output tokens on a DeepSWE subset This is pareto opt…
    @kimmonismusOx Alpha has now been fully tested against the DeepSWE set by @davis7 . Overall, it's performing more or less on par with GPT-5.6 Sol mid. But he makes a good point. If it really is GLM-5.3 Flash that's now performing at the level of 5.6 Sol mid and could run locally on a DGX Spark, I wouldn't just be satisfied, it would be an absolutely fantastic deal. 5.6 Sol mid running locally, with the only cost being power consumption, for all sorts of tasks would be incredibly great. This model 24/7 hermes local: game changer.
    @teortaxesTexalternative results I think 60-61% is realistic
    @thdxrlot of guesses on what ox alpha is but they are all wrong, kinda disappointed so just going to tell you ox alpha is a new kind of llm that recursively updates a persistent latent state instead of reasoning entirely through tokens this lets internal representations converge before anything is actually decoded those attractors generate shards that encode transformations between latent states rather than the states themselves at sufficient density these shards compose into metaparameters that dynamically alter the residual geometry of the model without changing its weights. we built it because there was one thing simply too large to fit inside the context window of any existing model your mom

    Sentiment

    Positive78.1%21.9%Negative

    Summary

    Many accounts welcomed Ox Alpha's near-GPT-5.6 Sol Mid performance on DeepSWE because it points to capable open models usable locally, while some questioned the evaluation subsets and production latency.

    Based on 33 sentiment-bearing replies from 32 accounts across 2 conversations.