Qwen introduces Qwen3.8-Omni-Flash for audio-video tasks and tool use
Qwen says the model can plan and execute tasks such as editing vlogs and translating short videos, with video input costs about 89% lower than Qwen3.5-Omni-Plus.
TLDR
Qwen says Qwen3.8-Omni-Flash combines audio-video understanding, reasoning and tool use, with a 1-million-token context. On OmniVideoBench, it claims higher accuracy in locating key moments in long videos while using 51.8% fewer tokens than static understanding. Qwen also announced it is open-sourcing Qwen-MM-Plugins and Qwen-Live Harness; its September 18 announcement listed the latter as “coming soon.”
Combined views
88
1 Source, first seen 1d ago
Qwen introduces Qwen3.8-Omni-Flash for audio-video tasks and tool use
Qwen says the model can plan and execute tasks such as editing vlogs and translating short videos, with video input costs about 89% lower than Qwen3.5-Omni-Plus.