Ox Alpha Scores 63 Percent on DeepSWE Subset
Posts discuss Ox Alpha results on DeepSWE benchmark and comparisons to other models.
A retweet from @theo shares @winkey_h noting Ox Alpha reached around 63 percent at 47K average output tokens on a DeepSWE subset, calling the result pareto optimal. A quote from @kimmonismus states that @davis7 fully tested Ox Alpha on the DeepSWE set, with overall performance described as more or less on par with GPT-5.6 Sol mid. The same post speculates the model may be GLM-5.3 Flash capable of local runs on hardware such as a DGX Spark. One generated headline in the packet refers to Ox Alpha as a stealth model from Opencode with 1M context length and multimodal support.
Combined views
547.3K
5 posts, first seen 3d ago



