• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Qwen launches Qwen3.8-Omni-Flash with audio-video reasoning and tool use

    Qwen says the model has a 1M-token context window and cuts video input costs by about 89% compared with Qwen3.5-Omni-Plus.

    QwenQW
    2 Sources, 20d ago, first seen 20d ago

    TLDR

    Qwen says Qwen3.8-Omni-Flash can reason over audio and video together, plan tasks and execute them with tools. Examples include editing vlogs, translating short videos and turning movies into recaps. The company claims about 89% lower video input costs than Qwen3.5-Omni-Plus. Qwen also announced it was open-sourcing Qwen-MM-Plugins and Qwen-Live Harness, though its September 18 announcement listed Live Harness as “coming soon.”

    Combined views

    261.1K

    2 Sources, first seen 20d ago

    Combined views

    261.1K

    2 Sources, first seen 20d ago

    3.5K likes
    3.5K likes
    153 comments
    1K saves
    356 reposts
    153 comments
    1K saves
    356 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    2 Sources

    智东西China AI News@Chinazhidx🚀Qwen just released Qwen3.8-Omni-Flash, its new native omni-modal model. It takes text, image, audio, and video inputs with a 1M-token context window. Across 29 benchmarks, it delivers an average improvement of 25%+ over Qwen3.5-Omni-Plus. The cost drop is even bigger: Audio input: 98%+ cheaper per hour Audio + video input: 93%+ cheaper per hour Qwen also released two companion frameworks: Qwen-MM-Plugins for multimodal perception and agent workflows, and Qwen-Live-Harness for continuous, real-time omni-modal interaction.20d
    Qwen@Alibaba_Qwen🚀 Meet Qwen3.8-Omni-Flash, Qwen's first omni-modal model built around agentic capabilities! Native audio-video understanding, reasoning, and tool use come together in one model: understand the content, plan the task, execute with tools, and deliver the result. Highlights: 🥳 - Audio-video intelligence that gets things done: jointly reason over what's seen and heard, and orchestrate tools across long workflows to auto-edit vlogs, translate short videos, and turn movies into recaps. - A major leap: approaching Gemini 3.8 Flash in audio-video capabilities; +19.5 points on average in agent performance across WildClawBench-MM & UniClawBench. - 1M-token context with agentic perception: actively explore long videos and locate key moments with higher accuracy, using 51.8% fewer tokens than static understanding on OmniVideoBench. Video input costs are reduced by about 89% compared with Qwen3.5-Omni-Plus, making long-form audio-video understanding and agentic workflows more affordable than ever. To help you build apps around Omni, we're also open-sourcing Qwen-MM-Plugins and Qwen-Live Harness! 🛠️ We can't wait to see what you build with Qwen3.8-Omni-Flash! 👀 - Blog: https://qwen.ai/blog?id=qwen3.8-omni-flash - Qwencloud: https://www.qwencloud.com/models/qwen3.8-omni-flash - Qwen Studio: https://chat.qwen.ai/ - API: https://www.alibabacloud.com/help/en/model-studio/qwen-omni - Qwen-MM-Plugins: https://github.com/QwenLM/Qwen-MM-Plugins - Qwen-Live Harness: coming soon https://github.com/QwenLM/Qwen-Live-Harness20d

    2 Sources

    智东西China AI News@Chinazhidx🚀Qwen just released Qwen3.8-Omni-Flash, its new native omni-modal model. It takes text, image, audio, and video inputs with a 1M-token context window. Across 29 benchmarks, it delivers an average improvement of 25%+ over Qwen3.5-Omni-Plus. The cost drop is even bigger: Audio input: 98%+ cheaper per hour Audio + video input: 93%+ cheaper per hour Qwen also released two companion frameworks: Qwen-MM-Plugins for multimodal perception and agent workflows, and Qwen-Live-Harness for continuous, real-time omni-modal interaction.20d
    Qwen@Alibaba_Qwen🚀 Meet Qwen3.8-Omni-Flash, Qwen's first omni-modal model built around agentic capabilities! Native audio-video understanding, reasoning, and tool use come together in one model: understand the content, plan the task, execute with tools, and deliver the result. Highlights: 🥳 - Audio-video intelligence that gets things done: jointly reason over what's seen and heard, and orchestrate tools across long workflows to auto-edit vlogs, translate short videos, and turn movies into recaps. - A major leap: approaching Gemini 3.8 Flash in audio-video capabilities; +19.5 points on average in agent performance across WildClawBench-MM & UniClawBench. - 1M-token context with agentic perception: actively explore long videos and locate key moments with higher accuracy, using 51.8% fewer tokens than static understanding on OmniVideoBench. Video input costs are reduced by about 89% compared with Qwen3.5-Omni-Plus, making long-form audio-video understanding and agentic workflows more affordable than ever. To help you build apps around Omni, we're also open-sourcing Qwen-MM-Plugins and Qwen-Live Harness! 🛠️ We can't wait to see what you build with Qwen3.8-Omni-Flash! 👀 - Blog: https://qwen.ai/blog?id=qwen3.8-omni-flash - Qwencloud: https://www.qwencloud.com/models/qwen3.8-omni-flash - Qwen Studio: https://chat.qwen.ai/ - API: https://www.alibabacloud.com/help/en/model-studio/qwen-omni - Qwen-MM-Plugins: https://github.com/QwenLM/Qwen-MM-Plugins - Qwen-Live Harness: coming soon https://github.com/QwenLM/Qwen-Live-Harness20d
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet