Alibaba launches Qwen3.8-Omni-Flash
The model combines audio-video understanding, reasoning and tool use, @thehypedotnews reports. In separate Qwen3.8-27B news, LM Studio announced day-one support for Inco AI's Splash inference engine.
TLDR
@thehypedotnews reports that Alibaba launched Qwen3.8-Omni-Flash, combining audio-video understanding, reasoning and tool use. Separately, Inco AI introduced Splash, an open-source inference engine built around Qwen3.8-27B and Apple silicon. Inco AI claims it runs that model at 144 tokens per second on an M5 Max MacBook Pro, with up to three times Ollama's decode speed. LM Studio says it partnered with Inco AI to make Splash available in LM Studio from day one.
Combined views
—
2 Sources, first seen ago