@arena @petergostev Thanks to OpenAI's Prism for writing LaTeX for me, Antigravity for running benchmark reports and monitoring 24/7, and Copilot Harness for unlimited access to GPT-4o/5 mini models (for students).
@arena @petergostev Qwen3.8-27B beating 100x larger models? that's the open-weight model finally eating the closed-weight lunch. efficiency > parameter bloat
@arena @petergostev The "up to 100x its size" part is wild—can’t wait to see how it holds up on agentic stuff, that’s where small models usually crumble for me.
Check out first impressions of Qwen3.8-27B with @petergostev on our YouTube. We tested it against models up to 100x its size (DeepSeek v4, Qwen 3.8 Max, Kimi K3, GLM 5.3, Grok 4.6, GPT-5.6, and Fable) on identical one-shot generation + agentic tasks. Scores coming soon.